@frank_uid: Recently studying Infra, vibed a Qwen3 inference engine implemented purely in C++/CUDA, with HuggingFace model parsing and benchmark totaling less than 2000 lines, completely dependency-free, compiled binary size only 1.2MB (Claude writes kernel too hardcore
Summary
FlashQwen is a minimal from-scratch C++/CUDA inference engine for Qwen3-8B with no external dependencies, supporting multi-turn streaming chat and benchmark mode, with a binary size of only 1.2MB.
View Cached Full Text
Cached at: 06/13/26, 02:46 PM
Recently Iβve been learning about Infra, and vibe-coded a Qwen3 inference engine. Itβs pure C++/CUDA, with HuggingFace model parsing and benchmark totaling less than 2000 lines, completely dependency-free. The compiled binary is only 1.2MB (Claude writes kernels like a beast π
) https://t.co/of3YkVZ7wm β # frankkk96/FlashQwen Source: https://github.com/frankkk96/FlashQwen # FlashQwen A minimal, from-scratch C++/CUDA inference engine for Qwen3-8B β no external libraries (no PyTorch, no cuBLAS, no tokenizers crate). Built for learning. Supports multi-turn streaming chat and a benchmark mode (TTFT / TPOT / tok/s). Chinese / δΈζζζ‘£: README.zh-CN.md β ## Usage ### Get the model git clone the model repo into a directory, then point --model at it: bash git lfs install git clone https://huggingface.co/Qwen/Qwen3-8B models/qwen3-8b The directory needs config.json, the *.safetensors shards, model.safetensors.index.json, vocab.json, and merges.txt. No offline conversion or repacking β FlashQwen reads the BF16 files directly and quantizes the matmul weights to INT8 in memory at load. ### Build bash cmake -B build -DCMAKE_BUILD_TYPE=Release cmake --build build -j Requires a CUDA toolkit (12.x recommended) and an NVIDIA GPU with β₯20 GB VRAM. The build targets sm_89 (RTX 4090 / Ada) by default; for a different GPU pass -DCMAKE_CUDA_ARCHITECTURES= (e.g. 90 for Hopper). ### Run --model is required and points at the model directory (e.g. models/qwen3-8b from above; an HF hub cache dir also works). --help lists supported models and scans the local hub cache for whatβs available. bash # interactive multi-turn chat (default mode). KV cache persists across turns. ./build/flashqwen --model models/qwen3-8b # chat with Qwen's recommended sampling ./build/flashqwen --model models/qwen3-8b --temperature 0.6 --top-p 0.95 --top-k 20 # benchmark (built-in fixed sweep over input lengths) ./build/flashqwen benchmark --model models/qwen3-8b ./build/flashqwen --help Modes: bare command (or chat) β interactive chat; benchmark β metrics. In-chat commands: /exit /quit /reset (clear history) /think on|off. Common flags: --max-ctx N (KV-cache size, default 4096), --temperature, --top-p, --top-k, --seed, --think (enable Qwen3 thinking mode). Supported models: any dense Qwen3 model (architecture Qwen3ForCausalLM): Qwen3-0.6B / 1.7B / 4B / 8B / 14B / 32B β dims are read from config.json. Not supported: Qwen3.5 (hybrid linear-attention + multimodal), Qwen3 MoE variants, and non-Qwen architectures. > First launch loads ~16 GB of weights from disk (slow on network storage); subsequent > runs hit the OS page cache and start in seconds. β ## Local benchmark Hardware: NVIDIA GeForce RTX 4090 (24 GB) Β· driver 580.76.05 (CUDA 13.0) Β· built with CUDA 12.8 (native sm_89) Model: Qwen3-8B β matmul weights quantized to INT8 (per-row scale), BF16 embeddings, FP32 activations; prefill on tensor cores (WMMA), decode attention via flash-decoding split-K, single-token decode replayed from a CUDA graph. Method: single stream (batch 1), greedy, fixed 128-token output (ignores EOS), 1 warmup run + median of 3 measured runs. The sweep is built in β just flashqwen benchmark --model . Swept over input length (synthetic prompts), output fixed at 128 tokens: | input tok | TTFT | TPOT | decode tok/s | output tok/s | peak tok/s | |β:|β:|β:|β:|β:|β:| | 16 | 69 ms | 9.6 ms | 104.2 | 98.7 | 105.1 | | 128 | 97 ms | 9.7 ms | 102.6 | 95.3 | 103.5 | | 512 | 0.40 s | 10.2 ms | 97.8 | 74.8 | 98.5 | | 1024 | 0.95 s | 10.9 ms | 91.9 | 54.6 | 92.6 | Metric definitions: - TTFT β time to first token = prefill of the whole prompt + sampling token #1. - TPOT β time per output token, averaged over the decode steps (after the first). - decode throughput β tokens/s during decode only. - output throughput β tokens/s including prefill (n_out / total_time). - peak output throughput β 1 / fastest single-token latency. Reading the numbers. Decode is memory-bound β each token reads the whole model β so the INT8 weights (~9 GB vs ~16 GB in BF16) put it around ~100 tok/s, and flash-decoding split-K keeps it nearly flat as context grows (TPOT 9.6 β 10.9 ms from 16 β 1024 tokens). Prefill runs on tensor cores (WMMA); the INT8 weights are dequantized to BF16 first, which is why TTFT is a touch higher than a pure-BF16 build. These are the numbers after the optimization study below β which starts from a 49.7 tok/s scalar baseline and shows where each gain came from. β ## Dependencies & code structure ### Dependencies β none third-party Only the C++ standard library, the CUDA Runtime (toolkit-bundled, not a 3rd-party library), and POSIX system calls. CMakeLists.txt has no target_link_libraries. With no dependencies to compile, a clean build (-j8) takes about 9 seconds, and the resulting binary is about 1.8 MB (1.6 MB stripped). The model weights are loaded from disk at runtime, so they are not part of the binary. | usually a library | here | |β|β| | nlohmann/json, rapidjson | hand-written src/json.hpp | | HF tokenizers / sentencepiece | hand-written byte-level BPE src/tokenizer.* | | safetensors C++ lib | hand-written src/safetensors.* (mmap + header parse) | | cuBLAS / CUTLASS | hand-written matmul (tensor-core WMMA for prefill, GEMV for decode) | | PyTorch (any DL framework) | not used | - C++ stdlib: β¦ - CUDA: , β no cuBLAS / cuDNN / cuRAND / Thrust. - POSIX: (mmap),, . ### Code structure & line count Each file has one job; `main.cpp` is a thin entry point that just parses arguments and dispatches. Roughly 1940 lines total. **Application layer** | file | role | lines | |---|---|---:| | `src/main.cpp` | entry point: argument parsing + dispatch | 69 | | `src/cli.cpp` / `.hpp` | architecture check + `--help` | 63 | | `src/chat.cpp` / `.hpp` | interactive multi-turn chat | 55 | | `src/benchmark.cpp` / `.hpp` | benchmark mode (input-length sweep) | 92 | | `src/generate.cpp` / `.hpp` | shared prefill + decode loop | 66 | | `src/sampler.cpp` / `.hpp` | token sampling (greedy / temp / top-k / top-p) | 50 | **Core engine** | file | role | lines | |---|---|---:| | `src/model.cu` / `.hpp` | weight loading (INT8 quant) + forward + KV cache + CUDA graph | 331 | | `src/kernels.cu` / `.cuh` | CUDA kernels (INT8 GEMV / WMMA matmul, attention, split-K, ...) | 431 | | `src/tokenizer.cpp` / `.hpp` | byte-level BPE encode/decode | 394 | | `src/safetensors.cpp` / `.hpp` | mmap + `.safetensors` header parse | 105 | | `src/json.hpp` | minimal JSON parser | 203 | | `src/config.hpp` | parse `config.json` | 45 | | `CMakeLists.txt` | build | 35 | Notably, the "usually-a-library" plumbing β tokenizer (394) + JSON (203) β is ~600 lines, nearly half the project. The actual neural network β CUDA kernels (431) + forward (331) β is ~760 lines, with the growth coming from the optimization stages below (INT8, split-K, CUDA graph) rather than the base Qwen3 transformer, which is small and regular. ## Optimization study The matmul/decode path was optimized in stages. Each stage is a **tagged commit on the `optimization-study` branch**, so any version can be checked out and re-measured. **Stage 0 (scalar matmul) is the baseline** that everything below is compared against. Measured on RTX 4090 (Qwen3-8B, BF16) β single stream, batch 1, greedy, output fixed at 128 tokens, swept over input length, median of 3 runs: | stage | branch tag | TTFT@128 | TTFT@1024 | TPOT@1024 | decode@16 (tok/s) | |---|---|---:|---:|---:|---:| | **0 Β· scalar matmul (baseline)** | `bench-0-scalar` | 1531 ms | 12585 ms | 40.1 ms | 49.7 | | 1 Β· tensor-core (WMMA) prefill | `bench-1-wmma` | 89 ms | 1234 ms | 40.1 ms | 49.7 | | 2 Β· BF16 KV cache | `bench-2-bf16kv` | 89 ms | 1233 ms | 40.0 ms | 49.7 | | 3 Β· warp attention (no barrier) | `bench-3-attn` | 83 ms | 908 ms | 39.4 ms | 49.8 | | 4 Β· vectorized GEMV decode | `bench-4-gemv` | 83 ms | 907 ms | 38.8 ms | 51.1 | | 5 Β· GPU argmax (greedy) | `bench-5-argmax` | 83 ms | 909 ms | 38.5 ms | 51.7 | | 6 Β· CUDA graph (decode) | `bench-6-cudagraph` | 83 ms | 909 ms | 38.1 ms | 52.9 | | 7 Β· flash-decoding split-K | `bench-7-splitk` | 83 ms | 910 ms | 18.9 ms | 56.9 | | 8 Β· INT8 weight quantization | `bench-8-int8` | 97 ms | 954 ms | 10.9 ms | **104.2** | vs the baseline: **~13β16Γ faster prefill**, and **~2Γ faster decode** that's nearly flat across context β decode@16 49.7 β 104.2 tok/s, TPOT@1024 40.1 β 10.9 ms. (INT8 trades a little prefill TTFT for the decode win.) Reproduce any stage: bash git checkout bench-1-wmma # or bench-0-scalar β¦ bench-4-gemv cmake -B build -DCMAKE_BUILD_TYPE=Release && cmake βbuild build -j8 ./build/flashqwen benchmark βmodel models/qwen3-8b `` Stage 0 β scalar matmul (baseline), bench-0-scalar. One warp computes one output element (a dot product) for every matmul, prefill and decode alike. No tensor cores. Prefill is brutal: a 1024-token prompt takes ~12.6 s to first token, because the prefill GEMMs are compute-bound and run with no tensor cores. Stage 1 β tensor-core (WMMA) prefill, bench-1-wmma. Prefill (many tokens) converts activations to BF16 and runs the matmul on tensor cores via WMMA (16Γ16Γ16, FP32 accumulate); decode (one token) keeps the GEMV. Prefill collapses ~14Γ (TTFT@1024 12585 β 1234 ms). Decode is untouched β itβs memory-bound, tensor cores donβt help. Stage 2 β BF16 KV cache, bench-2-bf16kv. The KV cache is stored BF16 instead of FP32. Speed is unchanged here β the attention kernel at this point is latency-bound (a serialized per-key __syncthreads reduction), not bandwidth-bound β but the cache halves (~1.2 GB β ~0.6 GB at 4096 ctx), i.e. ~2Γ the max context. The byte savings only pay off in stage 3. Stage 3 β warp attention, bench-3-attn. Attention is rewritten from βone block per (head,query) with a per-key block reductionβ to one warp per (head,query): each lane owns 4 dims, the per-key qΒ·k is a warp-shuffle reduction, online softmax in registers β no barriers. Prefill attention drops ~26 % (TTFT@1024 1233 β 908 ms). Decode TPOT is roughly flat: at batch 1 there are only 32 work-items (one per head), so decode attention is parallelism-bound, not per-key-cost-bound (a real decode win needs flash-decoding split-K). Stage 4 β vectorized GEMV decode, bench-4-gemv. The decode GEMV reads 8 elements per step with 16-byte vectorized loads instead of one scalar BF16 at a time. A modest gain (decode@16 49.8 β 51.1 tok/s) β the scalar version was already ~80 % of memory bandwidth. Stage 5 β GPU argmax for greedy, bench-5-argmax. Instead of copying all 151936 logits to the host every token and scanning them there, greedy decoding now runs an argmax reduction on the GPU and copies back a single int. Saves ~0.2β0.3 ms/token (decode@16 51.1 β 51.7 tok/s). Sampling (temperature > 0) still copies the full logits. Stage 6 β CUDA graph for decode, bench-6-cudagraph. A decode step launches ~430 small kernels (36 layers Γ ~12); some are so short theyβre launch-overhead-bound rather than hidden behind GPU work. The fixed single-token sequence is captured once into a CUDA graph and replayed each step (token id / position / past_len live in device buffers the kernels read, so the graph stays valid as context grows). Consistent ~0.4β0.5 ms/token (decode@16 51.7 β 52.9 tok/s, peak 55.2 β 56.5). Prefill (variable length) stays eager. Stage 7 β flash-decoding split-K, bench-7-splitk. At M=1 the warp attention (stage 3) had only 32 work-items (one per head), so decode attention was parallelism-bound and TPOT grew with context (18.9 ms@16 β 38.1 ms@1024). Now each headβs key range is split into ATTN_SPLITS (16) chunks computed by separate blocks β each a partial online-softmax β and a combine pass merges them, giving 16Γ the parallelism over the KV cache. TPOT@1024 collapses 38.1 β 18.9 ms (decode 26.3 β 53.0 tok/s, ~2Γ) and decode is now nearly flat across context (~18β19 ms everywhere), i.e. bound by the weight reads (GEMV), not attention. Prefill keeps the stage-3 kernel (already parallel enough). Stage 8 β INT8 weight quantization, bench-8-int8. Decode was bandwidth-bound at the bf16 ceiling (read all ~16 GB of weights per token β ~60 tok/s). The matmul weights (attention + MLP projections + lm_head) are quantized to INT8 with a per-output-row scale (symmetric, computed at load); embedding/norms stay as-is. Decode reads 1-byte weights and dequantizes in-kernel β decode@16 56.9 β 104.2 tok/s (~1.8Γ), TPOT@1024 18.9 β 10.9 ms, and the weight memory roughly halves (~16 β ~9 GB). Prefill dequantizes each weight to BF16 before the WMMA GEMM, which costs a little TTFT (e.g. @128 83 β 97 ms) β a fine trade since prefill runs once but decode runs every token. Output stays coherent (per-channel INT8 is mild for an 8B model). Still open: INT4 / grouped quantization for more decode headroom, activation-INT8 tensor cores to also speed prefill, and shared-memory tiling for the prefill WMMA.
Similar Articles
I built an open-source multi-agent SDLC harness that beats a cold Claude Code run on large repos, by learning the repo once. Real benchmarks (incl. where it loses) inside. [P]
AutoDev Studio is an open-source multi-agent AI coding harness that reduces costs by building a persistent repository knowledge base, outperforming cold-start Claude Code runs on large repos in most tasks.
Hugging Face releases The Stack v3 β largest open code dataset yet
Hugging Face released The Stack v3, the largest open-source code dataset to date, designed to enhance training for code-focused AI models.
[BIG DATASET RELEASE] - SupraLabs/reasoning-corpus-4K-5M-v1 - Train your tiny SLMs to think!
SupraLabs releases reasoning-corpus-4K-5M-v1, a 5M-sample reasoning dataset for training small language models (SLMs), featuring chain-of-thought traces and ChatML format, hosted on Hugging Face.
US and China just teamed up to back open-source AI
The US and China have jointly agreed to support open-source AI, marking a rare collaboration between the two rival nations in the AI field.
Curated 11 token-optimization tools into one installer β what am I missing?
A developer shares a curated installer that packages 11 token-optimization tools to help AI agents manage context and reduce token usage, asking the community for additional recommendations especially for multi-agent setups.