[Discussion] A 5KB pure x86-64 assembly engine for Gemma-2B (FP16, 4.6 tok/s on CPU)
Summary
A personal project presents PULSAR-ASM, a minimal LLM inference engine written entirely in flat x86-64 assembly (FASM) that runs Gemma-2B in FP16 with a 5.2 KB binary footprint, zero C/C++ runtime dependencies, and ~4.5-4.7 tokens/s on an older quad-core CPU — framed as a first-principles exploration for bare-metal and microcontroller LLM inference.
Similar Articles
Running Gemma 4 26B at 5 tokens/SEC on a 13-year-old Xeon with no GPU
A developer successfully runs Google's Gemma 4 26B mixture-of-experts model at about 5 tokens per second on a 13-year-old dual Xeon server without a GPU, using a modified version of ik_llama.cpp that works without AVX2 instructions.
You don't need a GPU to run gemma-4-26B-A4B
The author demonstrates that the Gemma-4-26B-A4B model runs efficiently on a CPU-only system using Koboldcpp, achieving 7 tokens per second on an old desktop, suggesting that powerful GPUs may not be necessary for local LLM inference.
I implemented a modern LLM in 700 lines of C
A developer created a minimal 700-line C implementation for running the Gemma 4 E2B LLM on CPUs, outperforming llama.cpp in speed.
Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
TurboFieldfare is an open-source Swift+Metal runtime that runs the Gemma 4 26B-A4B model on Apple Silicon Macs using only ~2GB of RAM by streaming experts from SSD, enabling inference on 8GB machines.
Gemma 4 E2B running in-browser at 255 tok/s using WebGPU kernels written by Fable 5
Gemma 4 is demonstrated running in-browser via WebGPU at 255 tokens per second, using kernels generated by Fable 5, showcasing efficient on-device inference.