@tunguz: Here is one big reason why this matters. Time spent on non-LLM inference tasks is only going to increase. However, tool…
Summary
A post highlights that 42% of time in modern agentic coding is spent on CPU-based tool use, which is inefficient and presents a major opportunity to redesign these tools for AI agents.
View Cached Full Text
Cached at: 05/24/26, 12:13 AM
Here is one big reason why this matters. Time spent on non-LLM inference tasks is only going to increase. However, tools that these AI system use are very inefficient and have been built from the ground up for CPU and human use. There is a huge untapped opportunity there to significantly improve those processes with AI agents in mind from the ground up.
SemiAnalysis (@SemiAnalysis_): FACT ALERT 🚨 : In modern agentic coding, 42% of the time is spent on CPU doing tool use such as editing files, running Bash scripts, running lints, etc. The economy of traditional cloud computing charges at $ per cpu core. In the economy of agents, the business model is $ per
Similar Articles
LLMs and performative productivity
A developer reflects on using AI agents and questions whether the apparent productivity gains are genuine or merely performative, noting that while tasks are completed faster, deep understanding and real value may be lost.
@JiaZhihao: I think this comes down to a classic systems tradeoff: generality vs. specialization. vLLM/SGLang cover a huge space of…
The article discusses the tradeoff between generality and specialization in AI inference engines, with vLLM and SGLang as examples, and notes that coding agents are reducing engineering costs for creating specialized engines.
@bentlegen: to make *actually* good software, you have to use it a lot but many people are spending more time in their coding agent…
Developer Ben Tlegen observes that building truly good software requires extensive personal use, but warns that many developers are now spending more time inside AI coding agents than in the actual applications they are creating.
Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics
The paper characterizes the resource and performance dynamics of LLM-based AI agents across tasks like question answering and coding, revealing bottlenecks and proposing optimizations that improve latency by up to 5.4×.
@levie: Agents already make up the majority of inference. This will quickly trend toward nearly all inference over the next yea…
The tweet asserts that AI agents now dominate inference traffic, with this trend expected to accelerate, leading to agents performing numerous tasks 24/7 and consuming vast amounts of tokens.