Tag
This paper argues that the optimal tokenizer vocabulary size is not fixed but depends on the deployment regime, such as batch size and inference volume. Through roofline analysis and experiments on A10G and A100 GPUs, it shows the lifecycle-optimal vocabulary can shift by up to 16x between on-device and datacenter serving, with minimal quality impact.