[Discussion] A 5KB pure x86-64 assembly engine for Gemma-2B (FP16, 4.6 tok/s on CPU)

Reddit r/LocalLLaMA Tools

Summary

A personal project presents PULSAR-ASM, a minimal LLM inference engine written entirely in flat x86-64 assembly (FASM) that runs Gemma-2B in FP16 with a 5.2 KB binary footprint, zero C/C++ runtime dependencies, and ~4.5-4.7 tokens/s on an older quad-core CPU — framed as a first-principles exploration for bare-metal and microcontroller LLM inference.

Hi everyone, Sharing a personal project exploring the minimal bare-metal footprint required to run an autoregressive LLM. Instead of relying on large runtimes or compiler abstractions, I wrote an inference engine for Gemma-2B entirely in flat x86-64 assembly (FASM): - **Binary footprint**: Total 5.2 KB flat machine code (`gemma_engine.bin` 3.7 KB + `mat_smp_f16c_gemm_avx2.bin` 1.5 KB). - **Execution**: Pure AVX2 + F16C with custom 4-thread SMP GEMM for prefill. Sustains ~18.5 GB/s memory bandwidth on commodity DDR4-2400. - **Decoding**: 4.5 ~ 4.7 tokens/s in FP16 on an older quad-core i5 desktop. - **Dependencies**: Zero C/C++ runtime, zero PyTorch. The Python harness only uses `ctypes` for `VirtualAlloc` and OS threads. This isn't meant to compete with feature-complete tools like llama.cpp. Rather, it's a first-principles exploration to see how cleanly a modern Transformer can be mapped to raw silicon, and to serve as a reference point for future micro-LLMs on resource-constrained microcontrollers (MCU/DSP). The repository is open source: - GitHub: https://github.com/tomtsai28/PULSAR-ASM - Architecture notes: https://github.com/tomtsai28/PULSAR-ASM/blob/main/doc/pulsar_asm_cpu_limit_retrospective.md Any code audits, observations, or thoughts on bare-metal inference are welcome.
Original Article

Similar Articles

You don't need a GPU to run gemma-4-26B-A4B

Reddit r/LocalLLaMA

The author demonstrates that the Gemma-4-26B-A4B model runs efficiently on a CPU-only system using Koboldcpp, achieving 7 tokens per second on an old desktop, suggesting that powerful GPUs may not be necessary for local LLM inference.