Tag
This paper presents an automated tensor scheduling approach for hybrid CPU-GPU LLM inference on consumer devices.
This paper introduces AgentStop, a lightweight supervisor that predicts and preemptively terminates local AI agent trajectories unlikely to succeed, reducing energy waste by 15-20% with minimal impact on task performance.