From Human Guidance to Autonomy: Agent Skill System for End-to-End LLM Deployment on Spatial NPUs

arXiv cs.LG Papers

Summary

This paper presents a two-stage methodology for end-to-end LLM deployment on spatial NPUs, progressing from human-guided development to an autonomous agent skill system. The system achieves speedups of 2.2x on prefill and 4.0x on decode for a reference model, and autonomously deploys eight additional LLMs on AMD XDNA 2 NPU with minimal human guidance.

arXiv:2606.07586v1 Announce Type: new Abstract: Spatial neural processing units (NPUs) provide an energy-efficient platform for edge LLM inference, but efficiently deploying an LLM end-to-end on such hardware remains labor-intensive. Although AI coding agents have begun to lower this cost, existing studies have largely focused on single-kernel optimization rather than end-to-end LLM deployment on resource-constrained spatial NPUs. We present a two-stage methodology, instantiated on the AMD XDNA 2 NPU, that progresses from human-guided development to agent autonomy. In the first stage, we develop a reference deployment of Llama-3.2-1B through human-guided agent assistance. The resulting implementation achieves a speedup of 2.2x on prefill and 4.0x on decode over the hand-optimized baseline, with the optimization trajectory and its lessons recorded as structured documentation throughout. In the second stage, we distill the documentation into an agent skill system consisting of eight phases, orchestrating the optimization and debugging skill sets, with numerical correctness strictly enforced at each phase. Using our agent skill system, we autonomously deploy eight additional decoder-only LLMs (Llama-3.2-3B, SmolLM2-1.7B, Qwen2.5-{0.5B, 1.5B, 3B}, Qwen3-{0.6B, 1.7B, 4B}) end-to-end on the AMD XDNA 2 NPU using the open-source compiler stack. To our knowledge, these models have not previously been deployed on AMD NPUs via any open-source software stack. Each deployment completes in 0.5-4 hours of agent wall time with almost no human guidance, and passes the numerical-correctness gates, demonstrating functional generalization to previously unencountered LLMs. Three of the eight match or exceed the sustained performance of our Llama-3.2-1B reference deployment, suggesting that the resulting implementations can be competitive without additional model-specific human engineering.
Original Article
View Cached Full Text

Cached at: 06/09/26, 08:46 AM

# Agent Skill System for End-to-End LLM Deployment on Spatial NPUs
Source: [https://arxiv.org/html/2606.07586](https://arxiv.org/html/2606.07586)
## From Human Guidance to Autonomy: Agent Skill System for End\-to\-End LLM Deployment on Spatial NPUs

Erwei WangZhiru ZhangSamuel Bayliss

###### Abstract

Spatial neural processing units \(NPUs\) provide an energy\-efficient platform for edge LLM inference, but efficiently deploying an LLM end\-to\-end on such hardware remains labor\-intensive\. Although AI coding agents have begun to lower this cost, existing studies have largely focused on single\-kernel optimization rather than end\-to\-end LLM deployment on resource\-constrained spatial NPUs\.

We present a two\-stage methodology, instantiated on the AMD XDNA™ 2 NPU, that progresses from human\-guided development to agent autonomy\. In the first stage, we develop a reference deployment of Llama\-3\.2\-1B through human\-guided agent assistance\. The resulting implementation achieves a speedup of 2\.2×\\timeson prefill and 4\.0×\\timeson decode over the hand\-optimized baseline, with the optimization trajectory and its lessons recorded as structured documentation throughout\. In the second stage, we distill the documentation into an agent skill system consisting of eight phases, orchestrating the optimization and debugging skill sets, with numerical correctness strictly enforced at each phase\.

Using our agent skill system, we autonomously deploy eight additional decoder\-only LLMs \(Llama\-3\.2\-3B, SmolLM2\-1\.7B, Qwen2\.5\-\{0\.5B, 1\.5B, 3B\}, Qwen3\-\{0\.6B, 1\.7B, 4B\}\) end\-to\-end on the AMD XDNA 2 NPU using the open\-source compiler stack\. To our knowledge, these models have not previously been deployed on AMD NPUs via any open\-source software stack\. Each deployment completes in 0\.5–4 hours of agent wall time with almost no human guidance, and passes the numerical\-correctness gates, demonstrating functional generalization to previously unencountered LLMs\. Three of the eight match or exceed the sustained performance of our Llama\-3\.2\-1B reference deployment, suggesting that the resulting implementations can be competitive without additional model\-specific human engineering\.

## IIntroduction

Large language models are increasingly deployed at the edge for lower latency, stronger privacy, and offline operation\. These settings impose strict power and thermal constraints, motivating the use of accelerators with higher energy efficiency than general\-purpose CPUs or GPUs\. Spatial neural processing units \(NPUs\) have emerged as a key class of accelerator for this regime\.

Spatial NPUs achieve efficiency by exposing explicit hardware management to the software stack\. Unlike processors with implicit cache hierarchies, spatial NPUs expose distributed on\-chip memories, explicit data\-movement scheduling, and tile\-level kernel placement directly to the programmer\. Deploying an LLM end\-to\-end onto NPUs with competitive performance typically takes substantial expert engineering to tackle all these challenges\.

Agentic systems for accelerator programming\[[3](https://arxiv.org/html/2606.07586#bib.bib8),[7](https://arxiv.org/html/2606.07586#bib.bib14),[5](https://arxiv.org/html/2606.07586#bib.bib3),[8](https://arxiv.org/html/2606.07586#bib.bib9),[13](https://arxiv.org/html/2606.07586#bib.bib10),[14](https://arxiv.org/html/2606.07586#bib.bib11),[12](https://arxiv.org/html/2606.07586#bib.bib6)\]reduce engineering effort by automating kernel generation\. However, no existing research primarily targets end\-to\-end LLM deployment on a resource\-constrained spatial NPU\. They also do not encode the full deployment workflow as composable Agent Skills\[[2](https://arxiv.org/html/2606.07586#bib.bib2)\]that can be invoked autonomously by a coding agent\. We address both gaps for end\-to\-end LLM deployment on the AMD XDNA™ 2 NPU\[[9](https://arxiv.org/html/2606.07586#bib.bib1)\]\.

Our approach consists of two stages, embodying a transitionfrom human guidance to agent autonomy\. In the first stage, we construct a reference Llama\-3\.2\-1B on NPU with human guidance and agent assistance, and document the optimization trajectory and its lessons throughout\. In the second stage, we distill the documentation into a reusable agent skill system that supports the autonomous deployment of additional LLMs\. The paper makes three novel contributions:

- •An end\-to\-end Llama\-3\.2\-1B deployment on the AMD XDNA 2 NPU \(§[III](https://arxiv.org/html/2606.07586#S3)\) that achieves a speedup of 2\.2×\\timeson prefill and 4\.0×\\timeson decode over the hand\-optimized baseline\. The optimization trajectory and its lessons are maintained as documents throughout, forming the foundation of the skill system\.
- •An agent skill system for end\-to\-end LLM deployment on NPUs \(§[IV](https://arxiv.org/html/2606.07586#S4)\), comprising an eight\-phase skill chain with strict numerical correctness gates, sets of optimization and debug skills that auto\-trigger on common patterns, and an independent evaluator agent that re\-runs every gate to prevent the generator from bypassing it\.
- •First open\-source end\-to\-end deployment of eight additional LLMs on the AMD XDNA 2 NPU \(§[V](https://arxiv.org/html/2606.07586#S5)\): Llama\-3\.2\-3B, SmolLM2\-1\.7B, Qwen2\.5\-\{0\.5B, 1\.5B, 3B\}, and Qwen3\-\{0\.6B, 1\.7B, 4B\}\. Each deploys autonomously in 0\.5–4 hours of agent wall time, with three of them matching or exceeding the sustained performance of our Llama\-3\.2\-1B reference\.

## IIBackground

Target hardware\.We target the AMD XDNA 2 NPU in Ryzen™ AI 300/400 Series processors\. Compute tiles are arranged in a 4×\\times8 array, each with a 64 KB L1 scratchpad\. Each column also includes a 512 KB memory tile \(L2\) shared by its four compute tiles, plus a shim tile bridging to host DDR\. Compute tiles communicate via configurable streaming interconnects and cascade connections to neighbors\.

Programming model\.We program the NPU through MLIR\-AIR\[[11](https://arxiv.org/html/2606.07586#bib.bib13)\], a platform\-agnostic compiler abstraction for spatial accelerators built on MLIR\. MLIR\-AIR defines the AIR dialect, which exposes the array’s spatial structure as loop nests\. AIR models spatial partitioning, temporal iteration, and inter\-tile communication explicitly, exposing the main deployment decisions for an AI agent to inspect and modify\.

IRON\[[6](https://arxiv.org/html/2606.07586#bib.bib7)\]is another programming abstraction that sits one layer closer to the hardware\. It exposes compute tiles, memory tiles, and shim tiles directly, with data movement wired through ObjectFifos and DMA tasks\. This gives expert programmers fine\-grained control over tiling, double buffering, data layout, and pipeline placement, well suited for hand\-tuned kernels\.

LLMs on AMD XDNA NPU\.AMD Ryzen AI Software\[[1](https://arxiv.org/html/2606.07586#bib.bib4)\]and FastFlowLM\[[4](https://arxiv.org/html/2606.07586#bib.bib5)\]deploy multiple LLMs on AMD XDNA NPUs, but they keep their NPU kernel implementation closed\-source and primarily target quantized LLMs\. Prior to our work, the only end\-to\-end LLM deployment with open\-source NPU kernels and compiler stack is the BF16 Llama\-3\.2\-1B example shipped with IRON, which we use as our baseline in §[III](https://arxiv.org/html/2606.07586#S3)\. MLIR\-AIR has no prior LLM deployment example\.

Agent skills\.Agent Skills\[[2](https://arxiv.org/html/2606.07586#bib.bib2)\]provide a portable format for encoding domain knowledge that a coding agent discovers and invokes autonomously\. Community skill collections such as Superpowers\[[10](https://arxiv.org/html/2606.07586#bib.bib15)\]apply this format to general software engineering\. To our knowledge, no existing skill set targets end\-to\-end LLM deployment on spatial NPUs\.

## IIIAgent\-Assisted Mapping of Llama\-3\.2\-1B

This section describes how we mapped a BF16 Llama\-3\.2\-1B end\-to\-end onto the AMD XDNA 2 NPU by working with an AI coding agent acting as our copilot\. The deployment outperforms the human\-engineered baseline by 2\.2×\\timeson prefill and 4\.0×\\timeson decode\. We also maintain a set of documents that log experiences, caveats, and the performance optimization trajectory, which becomes the reusable foundation for the skill system in §[IV](https://arxiv.org/html/2606.07586#S4)\.

### III\-ADocument\-Guided Workflow with a Coding Agent

Modern AI coding agents such as Claude Code can write code well, but on a long multi\-session project they need a human to plan and direct, and a persistent way to remember what has been done\. Development on the NPU makes this acute: tooling documentation is sparse, and opaque error messages surface frequently\. Without a running record of decisions and debugging, every new session starts blind, easily redoing past work or breaking past fixes\.

To avoid this, we pair a structured development plan with a set of Markdown documents alongside the code\. The plan has two parts: first reach end\-to\-end functional correctness for prefill and decode, then optimize each path’s performance separately\. Across both stages,plan\.mdcaptures the overall strategy,progress\.mdtracks the current phase, andissues\.mdaccumulates each non\-obvious bug, its root cause, and the workaround applied\. The performance stage addsperf\_opt\_traj\.md, which logs every optimization attempt with measured before/after results, alongside per\-kernel design notes \(gemm\.md,attention\.md,rope\.md, …\) that record each kernel’s design rationale, shape\-specific tuning, tile configurations, and known pitfalls\. The human\-engineered baseline with IRON serves as the performance reference at every iteration, both per\-kernel and end\-to\-end\.

Throughout, the coding agent is Claude Code paired with Claude Opus 4\.7 \(1M context\)\. The human directs and approves while the agent does the search, code edits, profiling, and tool runs\. In each session, the agent reads the relevant documents to recover state, proposes the next step against the plan, executes it under human supervision, and updates the documents as work progresses\. This documentation\-first discipline keeps the workflow durable across sessions\.

### III\-BChallenges and the Optimization Trajectory

![Refer to caption](https://arxiv.org/html/2606.07586v1/figures/prefill_optimization_trajectory_compact.png)Figure 1:Prefill optimization trajectory of Llama\-3\.2\-1B \(BF16, seq\_len=2048\) on the AMD XDNA 2 NPU\.Mapping the model end\-to\-end on NPU surfaced challenges in both correctness and performance\. We kept correctness in check throughout the trajectory with a CPU FP32 reference, gating every kernel and block change against ground truth\. Performance challenges fell into three categories, each addressed by specific steps in Figure[1](https://arxiv.org/html/2606.07586#S3.F1)\.

\(A\) Kernel efficiency at production shapes\.Reaching peak kernel performance requires several optimization techniques: vectorization \(Step 1\), shape\-specific tile tuning \(Step 4 GEMM\), fused kernel design \(Step 4 FlashAttention\), and scaling with tile\-array parallelism \(Step 8\)\. The main challenge is scaling kernels from validation shapes to production LLM shapes, where implementations must satisfy additional architectural and runtime constraints such as DMA\-channel availability, L1/L2 memory capacity, and BufferObject descriptor limits\. Cases that expose missing compiler coverage are reported upstream and resolved iteratively\.

\(B\) Reducing kernel dispatch overhead\.NPU kernel dispatch has non\-negligible overhead from the application, runtime, driver, firmware, and hardware layers\. This overhead can exceed the kernel’s actual execution time\. We merged consecutive kernels into single dispatches\. In prefill, we merged the eight\-kernel post\-attention block \(Step 5: output projection \+ SwiGLU FFN\) into one dispatch and the six\-kernel pre\-attention block \(Step 6: RMSNorm \+ QKV projections \+ RoPE\) into another, reducing per\-layer dispatches from 15 to 3\. The same merging applies to decode with GEMV variants\. This merging also saves memory copies of intermediate activations between the NPU and the host\.

\(C\) Host\-side optimization\.Host\-side overheads such as context setup, buffer management, host\-device data transfers, and layout transposes add up across many kernel calls\. We reused the XRT context across calls \(Step 2\) and recycled per\-layer buffer objects via zero\-copy mapping \(Step 3\)\. We skipped redundant host\-device transfers for buffers the NPU writes for static weights and intermediate activations \(Step 7\), and chose activation layouts that let consecutive kernels hand off without host\-side transpose \(Step 9\)\. Finally, prefill writes only the last token’s logits from the LM Head, removing a large NPU\-to\-host transfer of the full\-sequence logits tensor \(Step 10\)\.

As a result, on a 2048\-token sequence, prefill achieves a2\.2×\\timesspeedup over IRON \(commit 2b62dc7\), with a time\-to\-first\-token \(TTFT\) of 1\.3 s\. Decode achieves a4\.0×\\timesspeedup, reaching 10\.8 tokens/s \(TPS\)\. The trajectory and its lessons are logged and distilled into reusable agent skills in the next section\.

![Refer to caption](https://arxiv.org/html/2606.07586v1/figures/skill_system_overview.png)Figure 2:Overview of our skill system for end\-to\-end LLM deployment on NPU\. The deploy\-new\-llm orchestrator dispatches eight phases\. Phase 4 and 5 draw on optimization skills, and any phase can auto\-invoke a debug skill on known failures\. Phases 0\-6 each close with a numerical gate against the CPU FP32 reference and a human\-in\-the\-loop checkpoint\. Phase 7 spawns an independent evaluator that re\-audits the deployment from scratch\.Building on the experience of §[III](https://arxiv.org/html/2606.07586#S3), we distill the maintained document set into a self\-evolving agent skill system to automate the deployment of unencountered LLMs\. Figure[2](https://arxiv.org/html/2606.07586#S4.F2)shows the system and Table[I](https://arxiv.org/html/2606.07586#S4.T1)catalogs each skill\. Each maintained document maps to a skill component:plan\.mdandprogress\.mddefine the orchestrator and its phase skills, the per\-kernel design notes \(gemm\.md,attention\.md, …\) become a kernel registry used for Phase 1,perf\_opt\_traj\.mdbecomes the optimization skills applied in Phase 4 and Phase 5, andissues\.mdturns into the auto\-invoked debug skills\.

TABLE I:All skills in the system, with their type and purpose\.SkillTypePurposeDeploy new LLMOrchestratorScaffold workspace & dispatch phasesBuild CPU oraclePhase 0Decompose CPU FP32 oracle from HFKernel validationPhase 1Per\-kernel verification vs oracleSingle\-block validationPhase 2Single transformer block on NPUFull\-model validationPhase 3N\-layer cascade & final logits checkPrefill optimizationPhase 4Apply prefill perf patternsDecode optimizationPhase 5Apply decode perf patternsFinalize & learnPhase 6Runner integration & lesson harvestIndependent evaluatorPhase 7Recheck every gate on a fresh subagentMulti\-kernel mergingOptimizationMerge N kernels into one dispatchBuffer object reuseOptimizationLoad weights with zero\-copy mappingLayout alignmentOptimizationAlign layouts to skip host transposesRuntime failureDebugDiagnose kernel hang at runtimeBuffer object corruptionDebugDiagnose BO corruption runtime errorRouting congestionDebugDiagnose routing congestion hangThe eight\-phase decomposition mirrors the high\-level plan we followed in §[III](https://arxiv.org/html/2606.07586#S3)\. Phases 0–3 establish end\-to\-end functional correctness\. Phase 0 pulls the HuggingFace \(HF\) model and decomposes it into a kernel\-by\-kernel CPU FP32 reference verified against the HF output\. Phase 1 verifies each NPU kernel at the model’s required shapes\. Phase 2 wires those kernels into a single transformer block\. Phase 3 cascades to the full N\-layer model\. Phases 4 and 5 invoke optimization skills from a shared skill set to address prefill and decode bottlenecks\. Phase 6 integrates prefill and decode into a single inference runner, profiles the deployment to produce a performance report, and harvests new findings intolessons\.md\. Phase 7 spawns an independent evaluator agent without prior context to re\-audit the deployment\.

Each of Phases 0–6 closes with a strict numerical gate against the CPU FP32 reference, measuring absolute error, relative error, and Pearson correlation across the phase’s output\. The gate passes only if all three meet phase\-specific thresholds\. For Phases 4 and 5, the gate additionally requires that each optimization step measurably improves performance over the prior baseline\. If the gate fails, the agent invokes a matching debug skill on a known symptom, or asks the user if the failure is a new one\. If the gate passes, the agent stops at a human\-in\-the\-loop checkpoint where the user reviews the phase output and either approves or redirects\. Separating Phase 7’s evaluator from the deploying agent prevents reward hacking\[[15](https://arxiv.org/html/2606.07586#bib.bib12)\]\. The evaluator re\-runs every numerical gate from scratch and re\-profiles to verify that measured performance matches the profiling report submitted by the deploying agent\. A fail verdict blocks final acceptance\.

Two skill sets complement the phase pipeline: the optimization skills distilled from §[III\-B](https://arxiv.org/html/2606.07586#S3.SS2)patterns, and the debug skills distilled from high\-impact bug reports\. Each set currently contains three entries, and both are expected to grow gradually\. A new debugging fix or optimization pattern, once verified, becomes a new entry\. At the end of Phase 6, new experiences are also harvested intolessons\.mdand fed back into the skill system for subsequent deployments to inherit\.

## VEvaluation

TABLE II:End\-to\-end performance of the deployed LLMs on the AMD XDNA 2 NPU using our agent skill system\.ηscale\\eta\_\{\\text\{scale\}\}is the sustained performance on this model normalized to Llama\-3\.2\-1B’s\. Deploy time \(h\) is the agent’s wall clock for entire deployment\.ModelConfigPrefillDecodeDeploytime \(h\)𝑳\\boldsymbol\{L\}𝒅head\\boldsymbol\{d\_\{\\text\{head\}\}\}𝒉/𝒉𝒌​𝒗\\boldsymbol\{h/h\_\{kv\}\}Attn\.𝒅model\\boldsymbol\{d\_\{\\text\{model\}\}\}𝒅ffn\\boldsymbol\{d\_\{\\text\{ffn\}\}\}\|𝑽\|\\boldsymbol\{\|V\|\}QKV biasQK NormTTFT \(s\)𝜼scale\\boldsymbol\{\\eta\_\{\\text\{scale\}\}\}TPS𝜼scale\\boldsymbol\{\\eta\_\{\\text\{scale\}\}\}Llama\-3\.2\-1B†166432/8GQA20488192128k——1\.31\.0010\.81\.00—Llama\-3\.2\-3B2812824/8GQA30728192128k——3\.51\.074\.71\.270\.5∗SmolLM2\-1\.7B246432/32MHA2048819249k——2\.11\.027\.31\.220\.5∗Qwen2\.5\-0\.5B246414/2GQA8964864152k✓—0\.90\.568\.30\.284\.0∗Qwen2\.5\-1\.5B2812812/2GQA15368960152k✓—2\.60\.674\.90\.602\.5∗Qwen2\.5\-3B3612816/2GQA204811008152k✓—3\.90\.944\.21\.091\.8Qwen3\-0\.6B2812816/8GQA10243072152k—✓2\.30\.2410\.50\.481\.5∗Qwen3\-1\.7B2812816/8GQA20486144152k—✓2\.80\.686\.70\.941\.5∗Qwen3\-4B3612832/8GQA25609728152k—✓8\.00\.552\.60\.832\.1
†Reference deployment from §[III](https://arxiv.org/html/2606.07586#S3); the rest are deployed autonomously\.∗Estimated; the rest are measured from per\-phase timing logs\.

Experiment setup\.All deployments target the AMD XDNA 2 NPU in Ryzen AI 9 HX 370, programmed via MLIR\-AIR\. The coding agent is Claude Code paired with Claude Opus 4\.7 at max reasoning effort\. All evaluations use a sequence length of 2048\.

With our skill system, we autonomously deploy eight additional decoder\-only LLMs end\-to\-end\. To our knowledge, these are the first open\-source end\-to\-end LLM deployments on the AMD XDNA NPU, each completing in 0\.5–4 hours with almost no human intervention\. Human input is limited to picking among debug directions the agent proposes when it hits an unknown bug\.

Table[II](https://arxiv.org/html/2606.07586#S5.T2)reports the configurations and measured performance of the eight autonomously deployed LLMs alongside the Llama\-3\.2\-1B reference\. The set includes Multi\-Head Attention \(MHA\) and Grouped\-Query Attention \(GQA\), QKV bias \(Qwen2\.5 family\), per\-head Q/K normalization \(Qwen3 family\), and tensor shapes that are not always aligned to tile sizes\. All eight LLMs pass all correctness gates and are also manually verified\.

To assess how well the optimizations from our Llama\-3\.2\-1B reference transferred to the autonomous deployments of previously unencountered LLMs with different configurations, we reportηscale\\eta\_\{\\text\{scale\}\}for each model\.ηscale\\eta\_\{\\text\{scale\}\}is each new model’s sustained performance relative to Llama\-3\.2\-1B’s on the same hardware\. Performance here means achieved compute throughput for prefill \(compute\-bound\) and achieved memory bandwidth for decode \(memory\-bound\)\.ηscale\\eta\_\{\\text\{scale\}\}captures hardware utilization rather than absolute latency, so it is comparable across model sizes\.ηscale=1\\eta\_\{\\text\{scale\}\}=1means this model runs the hardware as efficiently as the reference,ηscale\>1\\eta\_\{\\text\{scale\}\}\>1means it exceeds the reference, andηscale<1\\eta\_\{\\text\{scale\}\}<1means lower efficiency\. Full formulas and calibration are in Appendix[A](https://arxiv.org/html/2606.07586#A1)\.

Across the eight deployed models,ηscale\\eta\_\{\\text\{scale\}\}ranges 0\.24–1\.07 on prefill and 0\.28–1\.27 on decode\. Three models \(Llama\-3\.2\-3B, SmolLM2\-1\.7B, Qwen2\.5\-3B\) reachηscale\\eta\_\{\\text\{scale\}\}of 0\.94–1\.27 across both prefill and decode\. All three share Llama\-3\.2\-1B’s basic transformer block \(GQA or MHA overdmodel≥2048d\_\{\\text\{model\}\}\\geq 2048withdffn≥8192d\_\{\\text\{ffn\}\}\\geq 8192\) and differ only in size; the existing kernels handle them as\-is, and their larger GEMMs amortize the per\-launch dispatch overhead identified in §[III\-B](https://arxiv.org/html/2606.07586#S3.SS2)more effectively than on the reference\. The remaining five models reachηscale\\eta\_\{\\text\{scale\}\}of 0\.24–0\.94, which we attribute to two compounding factors\. First, several have small per\-kernel work \(Qwen3\-0\.6B has the smallestdffnd\_\{\\text\{ffn\}\}in our set\), so the per\-launch dispatch overhead is not amortized\. Second, the Qwen families introduce ops that the current kernels do not yet fuse into the projection launches \(QKV bias for Qwen2\.5, per\-head Q/K normalization for Qwen3\), and shapes not aligned to tile sizes require padding that wastes useful compute\. Both are concrete targets for future kernel and agent skill iterations\.

## VIConclusion and Future Work

We presented a two\-stage methodology for end\-to\-end LLM deployment on the AMD XDNA 2 NPU, embodying a transition from human guidance to agent autonomy\. Stage one mapped Llama\-3\.2\-1B with a coding\-agent copilot, achieving 2\.2×\\timesprefill and 4\.0×\\timesdecode speedups over the hand\-optimized baseline, while documenting the trajectory throughout\. Stage two distilled the documentation into a reusable agent skill system, which we used to deliver the first open\-source end\-to\-end deployment of eight additional LLMs on the AMD NPU\. Each completes in 0\.5–4 hours, with three matching or exceeding the reference’s sustained performance\.

Several directions remain open\. New skills can target additional performance avenues such as kernel fusion, quantization, and dataflow optimization\. Architecture coverage can broaden to support sliding\-window attention, Mixture\-of\-Experts, and Multi\-head Latent Attention\. The methodology can also extend to other spatial accelerators and coding agents\.

## References

- \[1\]AMD\(2026\)AMD Ryzen AI Software\.Note:https://www\.amd\.com/en/products/software/ryzen\-ai\-software\.htmlAccessed: 2026\-04\-30Cited by:[§II](https://arxiv.org/html/2606.07586#S2.p4.1)\.
- \[2\]Anthropic\(2025\)Equipping Agents for the Real World with Agent Skills\.Note:Anthropic Engineering Blog,https://www\.anthropic\.com/engineering/equipping\-agents\-for\-the\-real\-world\-with\-agent\-skillsCited by:[§I](https://arxiv.org/html/2606.07586#S1.p3.1),[§II](https://arxiv.org/html/2606.07586#S2.p5.1)\.
- \[3\]K\. Cheng, L\. Wang, J\. Khuu, M\. Saroufim, W\. Chi, J\. Wang, and J\. Isaacson\(2026\)KernelAgent: Hardware\-Guided GPU Kernel Optimization via Multi\-Agent Orchestration\.Note:PyTorch Blog,https://pytorch\.org/blog/kernelagent\-hardware\-guided\-gpu\-kernel\-optimization\-via\-multi\-agent\-orchestration/Cited by:[§I](https://arxiv.org/html/2606.07586#S1.p3.1)\.
- \[4\]FastFlowLM\(2026\)Run LLMs on AMD Ryzen™ AI NPUs in Minutes\.Note:https://fastflowlm\.com/Accessed: 2026\-04\-30Cited by:[§II](https://arxiv.org/html/2606.07586#S2.p4.1)\.
- \[5\]C\. Hong, S\. Bhatia, A\. Cheung, and Y\. S\. Shao\(2025\)Autocomp: A Powerful and Portable Code Optimizer for Tensor Accelerators\.External Links:2505\.18574,[Link](https://arxiv.org/abs/2505.18574)Cited by:[§I](https://arxiv.org/html/2606.07586#S1.p3.1)\.
- \[6\]E\. Hunhoff, J\. Melber, K\. Denolf, A\. Bisca, S\. Bayliss, S\. Neuendorffer, J\. Fifield, J\. Lo, P\. Vasireddy, P\. James\-Roxby, and E\. Keller\(2025\)Efficiency, Expressivity, and Extensibility in a Close\-to\-Metal NPU Programming Interface\.In2025 IEEE 33rd Annual International Symposium on Field\-Programmable Custom Computing Machines \(FCCM\),pp\. 85–94\.Cited by:[§II](https://arxiv.org/html/2606.07586#S2.p3.1)\.
- \[7\]S\. Kalade and G\. Schelle\(2025\)NPUEval: Optimizing NPU Kernels with LLMs and Open Source Compilers\.External Links:2507\.14403,[Link](https://arxiv.org/abs/2507.14403)Cited by:[§I](https://arxiv.org/html/2606.07586#S1.p3.1)\.
- \[8\]G\. Liao, H\. Qin, Y\. Wang, A\. Golden, M\. Kuchnik, Y\. Yetim, J\. J\. Ang, C\. Fu, Y\. He, S\. Hsia, Z\. Jiang, D\. Li, U\. Pashkevich, V\. Puvvada, F\. Shi, M\. Steiner, R\. Xiao, N\. Yan, X\. Yu, Z\. Fang, R\. Levenstein, K\. Ho, H\. Zhu, A\. Hammond, R\. Li, A\. Mathews, K\. Gondkar, A\. Zainul\-Abedin, K\. Singh, H\. Yu, W\. Chi, B\. Huang, S\. Zhang, N\. Weller, Z\. Marine, W\. Cook, C\. Wu, and G\. Liu\(2026\)KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta\.External Links:2512\.23236,[Link](https://arxiv.org/abs/2512.23236)Cited by:[§I](https://arxiv.org/html/2606.07586#S1.p3.1)\.
- \[9\]A\. Rico, S\. Pareek, J\. Cabezas, D\. Clarke, B\. Ozgul, F\. Barat, Y\. Fu, S\. Münz, D\. Stuart, P\. Schlangen, P\. Duarte, S\. Date, I\. Paul, J\. Weng, S\. Santan, V\. Kathail, A\. Sirasao, and J\. Noguera\(2024\)AMD XDNA NPU in Ryzen AI Processors\.IEEE Micro44\(6\),pp\. 73–82\.External Links:[Document](https://dx.doi.org/10.1109/MM.2024.3423692)Cited by:[§I](https://arxiv.org/html/2606.07586#S1.p3.1)\.
- \[10\]J\. Vincent\(2026\)Superpowers: An Agentic Skills Framework & Software Development Methodology that Works\.Note:https://github\.com/obra/superpowersAccessed: 2026\-04\-30Cited by:[§II](https://arxiv.org/html/2606.07586#S2.p5.1)\.
- \[11\]E\. Wang, S\. Bayliss, A\. Bisca, Z\. Blair, S\. Chowdhary, K\. Denolf, J\. Fifield, B\. Freiberger, E\. Hunhoff, P\. James\-Roxby, J\. Lo, J\. Melber, S\. Neuendorffer, E\. Richter, A\. Rosti, J\. Setoain, G\. Singh, E\. Taka, P\. Vasireddy, Z\. Yu, N\. Zhang, and J\. Zhuang\(2025\)From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR\-AIR\.ACM Transactions on Reconfigurable Technology and Systems\.Cited by:[§II](https://arxiv.org/html/2606.07586#S2.p2.1)\.
- \[12\]J\. Wang, V\. Joshi, S\. Majumder, X\. Chao, B\. Ding, Z\. Liu, P\. P\. Brahma, D\. Li, Z\. Liu, and E\. Barsoum\(2025\)Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks\.External Links:2507\.23194,[Link](https://arxiv.org/abs/2507.23194)Cited by:[§I](https://arxiv.org/html/2606.07586#S1.p3.1)\.
- \[13\]A\. Wei, T\. Sun, Y\. Seenichamy, H\. Song, A\. Ouyang, A\. Mirhoseini, K\. Wang, and A\. Aiken\(2025\)Astra: A Multi\-Agent System for GPU Kernel Performance Optimization\.External Links:2509\.07506,[Link](https://arxiv.org/abs/2509.07506)Cited by:[§I](https://arxiv.org/html/2606.07586#S1.p3.1)\.
- \[14\]G\. Zhang, S\. Zhu, A\. Wei, Z\. Song, A\. Nie, Z\. Jia, N\. Vijaykumar, Y\. Wang, and K\. Olukotun\(2026\)AccelOpt: A Self\-Improving LLM Agentic System for AI Accelerator Kernel Optimization\.External Links:2511\.15915,[Link](https://arxiv.org/abs/2511.15915)Cited by:[§I](https://arxiv.org/html/2606.07586#S1.p3.1)\.
- \[15\]L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. Xing, H\. Zhang, J\. Gonzalez, and I\. Stoica\(2023\)Judging LLM\-as\-a\-Judge with MT\-Bench and Chatbot Arena\.InAdvances in Neural Information Processing Systems,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Vol\.36,pp\. 46595–46623\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/91f18a1287b398d378ef22505bf41832-Paper-Datasets_and_Benchmarks.pdf)Cited by:[§IV](https://arxiv.org/html/2606.07586#S4.p3.1)\.

## Appendix ARoofline derivation ofηscale\\eta\_\{\\text\{scale\}\}

ηscale\\eta\_\{\\text\{scale\}\}measures how efficiently a deployed model uses the NPU relative to our Llama\-3\.2\-1B reference\. Higher is better;ηscale\>1\\eta\_\{\\text\{scale\}\}\>1means that the deployed model exceeds the Llama\-3\.2\-1B reference\. We define

ηscale=Wmodel/tmodelWref/tref\\eta\_\{\\text\{scale\}\}=\\frac\{W\_\{\\text\{model\}\}/t\_\{\\text\{model\}\}\}\{W\_\{\\text\{ref\}\}/t\_\{\\text\{ref\}\}\}\(1\)whereWWis the dominant work—FLOPs \(FF\) for prefill, bytes \(BB\) for decode—andttis the measured time\. The numerator is sustained performance \(achieved compute throughput for prefill, achieved memory bandwidth for decode\); the denominator normalizes it to the reference\. We estimateFFandBBbelow from the model configuration\.

Prefill \(compute\-bound\):FprefilllayerF^\{\\text\{layer\}\}\_\{\\text\{prefill\}\}\.We estimate prefill FLOPs by counting the seven linear projections \(QKV, O, and SwiGLU FFN’s gate/up/down\) and FlashAttention; embeddings, RoPE, and norms are negligible and dropped:

Fprefilllayer​\(S\)=2​S​dmodel​\(2​dmodel\+2​hk​v​dhead\+3​dffn\)\+2​S2​dmodelF^\{\\text\{layer\}\}\_\{\\text\{prefill\}\}\(S\)=2S\\,d\_\{\\text\{model\}\}\(2d\_\{\\text\{model\}\}\+2h\_\{kv\}d\_\{\\text\{head\}\}\+3d\_\{\\text\{ffn\}\}\)\+2S^\{2\}d\_\{\\text\{model\}\}\(2\)The bracketed term sums the seven projections;2​S2​dmodel2S^\{2\}d\_\{\\text\{model\}\}is causal attention\. GQA is captured throughhk​vh\_\{kv\}\(MHA is the special casehk​v=hh\_\{kv\}=h\)\.

Decode \(memory\-bound\):BdecodelayerB^\{\\text\{layer\}\}\_\{\\text\{decode\}\}\.At each generated token, the NPU loads the layer’s weights from DRAM \(re\-loaded per token at batch=1\) and reads the KV cache for theSSpast tokens\. Other transfers—activation vectors, norm parameters, the new K/V written back—are orders of magnitude smaller and we drop them\. With weights and KV cache both in BF16:

Bdecodelayer​\(S\)=\\displaystyle B^\{\\text\{layer\}\}\_\{\\text\{decode\}\}\(S\)=\{\}2​dmodel​\(2​dmodel\+2​hk​v​dhead\+3​dffn\)\\displaystyle 2\\,d\_\{\\text\{model\}\}\(2d\_\{\\text\{model\}\}\+2h\_\{kv\}d\_\{\\text\{head\}\}\+3d\_\{\\text\{ffn\}\}\)\(3\)\+4​S​hk​v​dhead\\displaystyle\{\}\+4\\,S\\,h\_\{kv\}d\_\{\\text\{head\}\}The first term is the weight bytes \(parameters from Eq\.[2](https://arxiv.org/html/2606.07586#A1.E2)’s bracket,×\\times2 for BF16\); the second is the KV cache \(K and V, eachS×hk​v×dheadS\\times h\_\{kv\}\\times d\_\{\\text\{head\}\}BF16 elements\)\.

Similar Articles

Bayesian-Agent: Posterior-Guided Skill Evolution for LLM Agent Harnesses

Hugging Face Daily Papers

Bayesian-Agent presents a framework that treats reusable skills and SOPs as hypotheses, using Bayesian inference to guide agent behavior and improve task performance through posterior-guided harness optimization. It achieves significant improvements on multiple benchmarks with deepseek-v4-flash.

Towards Scalable Customization and Deployment of Multi-Agent Systems for Enterprise Applications

arXiv cs.CL

This paper proposes a unified framework for customizing and deploying LLM-based multi-agent systems in enterprise settings, combining model customization through continual pretraining, fine-tuning, and preference optimization with inference optimization using speculative decoding and FP8 quantization. It achieves 4.48x throughput speedup while maintaining performance on enterprise workloads.