Tag
A personal project presents PULSAR-ASM, a minimal LLM inference engine written entirely in flat x86-64 assembly (FASM) that runs Gemma-2B in FP16 with a 5.2 KB binary footprint, zero C/C++ runtime dependencies, and ~4.5-4.7 tokens/s on an older quad-core CPU — framed as a first-principles exploration for bare-metal and microcontroller LLM inference.