Tag
Inspired by CS336, a static performance model for LLM inference provides analytical bounds for VRAM, time-to-first-token, and throughput, covering various configurations and calibrated against public benchmarks.