Tag
LeanStream is a streaming speculate-and-refine framework that enables efficient on-device LLM inference by progressively refining computation and I/O operations, reducing memory usage and improving throughput.
This paper explores how prompt properties like cognitive load and phrasing pattern influence energy usage in on-device LLM inference, showing that cognitive load affects energy per token while phrasing impacts token usage, highlighting the need for model-aware prompt design for energy efficiency.