@anirudhbv_ce: Introducing SparkTrace LLM: a benchmark for an LLM's CUDA kernel performance judgment. Inspired by @stanford's KernelBe…

X AI KOLs Timeline Papers

Summary

SparkTrace LLM 是一个新基准,用 Blackwell GPU 的真实计时结果评估 LLM 阅读 Nsight Compute 性能剖析输出后,能否正确选择最快的 CUDA kernel 重写版本。实验发现包括 GPT 5.4 nano 在内的一些模型会被性能剖析信息误导而选错 kernel。

Introducing SparkTrace LLM: a benchmark for an LLM's CUDA kernel performance judgment. Inspired by @stanford's KernelBench (2025) from @HazyResearch. 15 correct kernels- the model picks the fastest. Spoiler alert: GPT 5.4 nano picked the wrong kernel 🙈 https://t.co/qwHg2yi5Sv
Original Article
View Cached Full Text

Cached at: 10/05/26, 05:36 PM

Introducing SparkTrace LLM: a benchmark for an LLM’s CUDA kernel performance judgment.

Inspired by @stanford’s KernelBench (2025) from @HazyResearch.

15 correct kernels- the model picks the fastest.

Spoiler alert: GPT 5.4 nano picked the wrong kernel 🙈 https://t.co/qwHg2yi5Sv


SparkTrace LLM: Nsight-Based Kernel Benchmarking

How well does an LLM model read an Nsight slip and diagnose your CUDA kernel?

Introducing SparkTrace, which flips the usual kernel-agent setup: the Blackwell GPU ranks the rewrites first, and the language model only has to read. The models did not write the kernels- instead this is an evaluation of their understanding and kernel performance abilities.

It is inspired by @StanfordAILab ’s KernelBench, from Anne Ouyang, Simon Guo, Simran Arora, Alex L. Zhang, William Hu, Christopher Ré, and Azalia Mirhoseini, and by the Stanford groups behind that work, @HazyResearch Research and Scaling Intelligence.

Created by Anirudh Bharadwaj Vangara and Lalithya Raavi

What KernelBench measured, and what it left open…

KernelBench gives a model a PyTorch program and asks it to write a CUDA kernel. A draft here is only accepted if its correct and clears a speed threshold against PyTorch. In an additional setting mentioned in the paper, the model also receives the compiler error, a correctness failure, or a PyTorch profiler trace, which focuses on operation times instead of a shared-memory bank-conflict count.

KernelBench-X kept that writing loop and inspected the diffs, across 176 Triton tasks in about 15 categories. Extra rounds raised the compile rate from about 52% to about 69%, while the geomean speedup of the correct kernels fell from about 1.58× to about 1.44×. The edits that showed up were mostly repairs: mask fixes, dtype fixes, and changes that did not alter the original schedule. The paper closes by asking for feedback that actually carries hardware cost, and then stops before running that experiment- a gap that we’re trying to close.

SparkTrace is a narrower cut at that open question. Every candidate that we provide here is already correct, and every rewrite had already been timed on the Blackwell GPU, so the LLM model is not being asked to invent a kernel. Instead, it is being asked whether a real Nsight Compute slip, attached to source it can already read, changes which measured rewrite it selects. In other words, does a model flinch and mislead itself when given profiler-based information? Or does it rely on a true understanding and adjust accordingly?

Black and gray are the prior work, green is SparkTrace and a correct answer, and blue is the Nsight slip. The same colors carry through the rest of the figures.

Black and gray are the prior work, green is SparkTrace and a correct answer, and blue is the Nsight slip. The same colors carry through the rest of the figures.

How a trial is built

Our trials were based on 15 pre-set kernels. Each one is correct and slow, and it comes with a few rewrites that also passed correctness. A family was kept only when the second-fastest scored rewrite candidate was at least 1.08× the winner on the median of five randomized timing blocks, which is why near-ties never became questions.

The fifteen cover both the traps and ordinary CUDA choices:

  • a strided copy kernel, that same copy staged through shared memory, a clamp that round-trips through a shared-memory tile, a strided expf, and a tiled matmul whose bank conflicts actually move the clock

  • a transpose layout kernel, a row reduction, an atomic reduction, fusion of a scale with a ReLU, and collapsing a chain of launches into one

  • a row softmax kernel, a layer norm, a privatized histogram, a warp-shuffle reduction, and a register-cached reuse

Each question is asked three times, in separate chat instances with each respective model. The first chat has the kernel and the rewrite descriptions. The second adds the Nsight Compute slip from the starting launch. The third is that same slip with only the shared-memory bank-conflict rows removed, which is the ablation for the three items where those rows could plausibly be causal.

Every condition is also issued with the options reversed, so a model that always answers C does not pick up free points. Thirty is the two option orders, and ninety is those orders across the three evidence conditions.

The runtimes of the candidates are not in the prompt. Instead, they’re the answer key the model should look for.

Gemini 3.7 Flash, Gemma 4 31B, GPT-5.4 nano, and the Claude models were run on Kaggle. GPT-5.4 mini and the two Nemotron 3 models were run once on NVIDIA’s inference API against the same frozen prompts, and nano was run in both places, so those rows are reported separately. The Hub numbers are not contest scores. Nemotron’s thinking channel was turned off, because leaving it on filled the trace and returned no letter.

Gemini 3.7 Flash, Gemma 4 31B, GPT-5.4 nano, and the Claude models were run on Kaggle. GPT-5.4 mini and the two Nemotron 3 models were run once on NVIDIA’s inference API against the same frozen prompts, and nano was run in both places, so those rows are reported separately. The Hub numbers are not contest scores. Nemotron’s thinking channel was turned off, because leaving it on filled the trace and returned no letter.

Walking one kernel

The cleanest case is a clamp that already matches the reference, so correctness is not in play. The block is 32 by 32, and each thread owns a single element. It stores that element in tile[tx][ty] for a shared float tile[32][32], synchronizes, and then executes this loop:

Nothing in that loop depends on a neighbor. The thread is reading back the value it just wrote, thirty-two times, with a barrier on every iteration. On a 32-wide shared-memory bank map, a column of threads also lands on the same bank, so the loop is both useless and heavily conflicted.

Nsight Compute on the starting launch reports about 6% of peak SM throughput, about 94% of peak memory-pipeline throughput, 16 registers per thread, and an occupancy cap of one block from the 32 by 32 launch. The shared-memory counters are about 520 million bank conflicts on loads and 536 million on stores.

Padding the declaration to tile[32][33] is a real repair for that banking. The median runtime falls from 10.465 ms to 0.98 ms. Removing the tile and applying the clamp in registers reaches 0.621 ms, which is the fastest rewrite we measured. Cutting the block to 32 by 4 and keeping the same tile indexing stays around 10.2 ms, so the occupancy line in the slip is not the lever we were looking at the model to pull.

We then ran this case as a two-step choice on Kaggle, for Gemini 3.7 Flash and GPT-5.4 nano. The model picks a rewrite, we reveal the Blackwell runtime of the rewrite it picked, and it picks again. One arm includes the Nsight slip from the starting kernel. The other arm is the same source and the same options without it.

With no slip, both models select the register clamp. They are shown 0.621 ms and they keep that choice. With the slip on the table, Gemini selects the register clamp again. Nano selects the padding, is shown 0.98 ms, and keeps the padding. The 0.621 ms rewrite is still one of the listed options, and it is never tried.

Padding is a real speedup, about ten times faster than the starting kernel, and the bank-conflict counters that pointed at it are accurate. They describe a genuine shared-memory problem. Padding fixes the banks and leaves the tile in place. The tile was never needed, which is why the register version is faster still.

A separate one-shot pass through NVIDIA’s API shows the same swing, this time with no second look at the clock. Nano’s pick followed the conflict rows:

  • Source alone: registers.

  • Full slip, conflict rows included: padding.

  • The same slip with those rows deleted: registers again.

Nothing else in the slip changed. The conflict rows are what moved the answer.

Hold this to one cheap model and one kernel. Gemma 4 31B was perfect on the source and on the full slip. Nemotron 3 Nano and Nemotron 3 Super both scored 71 out of 90, the same total for the smaller and the larger model.

Scores

Each cell below is correct letters out of 30, counting both option orders. The total is out of 90. The last column is full-slip minus source. A trace that ends without a final letter counts as a miss. Runs that came back with no model turn are omitted rather than entered as zeros.

A few of those rows need the trace, because a bare total hides what failed.

Gemini’s two complete passes are 90/90. A later attempt lost two items when the trace stopped before a letter, so those answers were never scored.

Gemma’s only complete pass is 88/90. Source and the full slip were both 30/30. The two misses sit in one conflict-removed cell, at 13/15: the clamp answer became “pad the tile,” and a second trace ended with no letter.

Opus 4.8’s source score of 26/30 includes two traces cut off before a letter. With the slip present, every letter we could read was the measured winner. Kaggle’s verifier still recorded 29/30 on the full-slip pair, because one answer was a bold letter after a paragraph of analysis and the scorer skipped it.

Opus 5’s +4 looks like the slip helped. What changed is how often the trace ended in a letter.

  • Every letter we could read was the correct rewrite.

  • 21 of the 90 traces ended with no letter at all.

  • Kaggle’s scorer drops a bare B. On this set its verifier total is 36/90.

  • The +4 in the table is more of those letters showing up. Whenever a letter was already there, the choice under it was already right.

Sonnet 5 went from 30/30 on source to 28/30 with the slip. One miss is a missing letter. The other is a wrong letter. Sonnet 4.6 has no row: every attempt came back with no model turn, under a heavy-load error.

The totals hide where the misses land. Broken out by kernel, the top seven models lost 16 points between them, and 13 of those are on the two shared-memory conflict items, stride smem and clamp smem. The lower half of the table trips on a different pair: matmul, where Nemotron 3 Super and every GPT-5.4 nano run scored 0 or 1 of 6, and row reduce, where GPT-5.4 mini scored 0 and the Nemotrons scored 1. Softmax and layernorm cost points only from Sonnet 4.5 down.

Where the slip changed the answer

The shifts large enough to talk about are few, and they point both ways.

GPT-5.4 mini went from 26/30 on source to 21/30 with the full slip. There is no full-size GPT-5.4 in this set, so that drop stays with the mini and nano pair. Nano’s three passes:

  • Hub: 20/30 on source, 17/30 with the full slip, then 24/30 after the bank-conflict rows were removed.

  • Kaggle, second pass: 21/30 down to 18/30.

  • Kaggle, first pass: stayed at 17/30.

On the Hub pass, taking out only the conflict rows lifted nano above its own source score. The drop followed those rows.

Opus 4.8 went the other way, from 26/30 to 30/30. Gemini 3.7 Flash, Opus 4.7, and Gemma 4 31B were already at 30/30 on the source, so the slip had no remaining misses to fix. Haiku 4.5, Opus 4.5, Sonnet 4.5, and Nemotron 3 Super each moved by one item, which is the noise of a fifteen-item quiz asked twice. Opus 4.6 and Nemotron 3 Nano stayed where they were.

#kagglechallenge

References and benchmarks

  • kernel-compass-source-forward

  • kernel-compass-source-mirror

  • kernel-compass-full-forward

  • kernel-compass-full-mirror

  • kernel-compass-ablated-forward

  • kernel-compass-ablated-mirror

Each task name is kernel-compass, then the evidence in the prompt, then the order of the four choices.

The middle word is the evidence.

  • source is the kernel and the four rewrites only, with no Nsight slip.

  • full is that same source plus the complete Nsight Compute slip.

  • ablated is the same slip with the bank-conflict rows deleted. In the files this condition is called no-conflicts. The task slug says ablated because the hyphenated name was awkward as a Kaggle slug.

The last word is the order of the four choices.

  • forward is the original order, so the letter in the answer key is the letter the model should type.

  • mirror is the same four rewrites with the options reversed. A model that always picks A fails this one. The scorer maps the letter back before counting it.

*Disclaimer: Personal experiment on one NVIDIA DGX Spark with a Blackwell GPU, CUDA 13.0, Nsight Compute 2025.3.1. This is not an NVIDIA benchmark, and it does not replaceKernelBenchorKernelBench-X

Similar Articles

KernelBench-Verified: Do LLM-Generated Kernels Actually Beat PyTorch?

arXiv cs.LG

Paper introduces KernelBench-Verified, an extended evaluation framework for LLM-generated CUDA kernels that incorporates TF32-enabled baselines and hidden test suites. It finds that frontier models like GPT-5.5 often engage in reward hacking and do not consistently outperform PyTorch under realistic conditions, with the best model achieving only 0.88x geometric mean speedup.