mobile-computing

Tag

Cards List
#mobile-computing

LeanStream: A Speculate-and-Refine Streaming Framework for Efficient on-Device LLM Inference

arXiv cs.LG · 3d ago Cached

LeanStream is a streaming speculate-and-refine framework that enables efficient on-device LLM inference by progressively refining computation and I/O operations, reducing memory usage and improving throughput.

0 favorites 0 likes
#mobile-computing

How Do Prompt Variations Affect Energy Consumption in On-Device LLMs?

arXiv cs.CL · 4d ago Cached

This paper explores how prompt properties like cognitive load and phrasing pattern influence energy usage in on-device LLM inference, showing that cognitive load affects energy per token while phrasing impacts token usage, highlighting the need for model-aware prompt design for energy efficiency.

0 favorites 0 likes
← Back to home

Submit Feedback