Tag
Perplexity CEO Aravind Srinivas points out that the value in the AI industry is shifting from average users to heavy users. These users can consume large amounts of computing power. For example, a Meta engineer spent nearly $10 million a year on programming tools, and on Perplexity Computer, there are users spending over $10,000 per month running agent loops.
A multi-tweet analysis of ~15 agentic-loop papers concludes that the verifier, not the model, is the key predictor of success, with examples showing that robust, non-gamable checks (e.g., compilers, tests, verifiable rewards) dramatically improve performance, while failures stem from lack of such verifiers or gaming vulnerabilities.
Anthropic is reportedly paying less per GPU than Google for SpaceX compute, with Google paying $920m for 110k GPUs compared to Anthropic's $1.25b for 220k GPUs plus additional capacity, highlighting a significant cost discrepancy.
Nvidia's VP states compute costs now exceed employee costs for his team; Uber confirms by exhausting its 2026 AI coding budget by April due to high token costs.
MiniMax's new m3 model achieves the same score as Opus 4.7 on terminal-bench 2.1 while using 1/20th the compute and cost, attributed to their novel MiniMax Sparse Attention architecture.
A commentary questioning why users cannot run Gemini and Claude Code locally on their own GPUs, implying compute cost constraints are limiting access to these AI models.
A user shares their experience spending ¥3,000 worth of tokens to create an AI video, suggesting that creativity, rather than compute power, is the true bottleneck.
This position paper argues that AI agents are a production technology rather than a labor input, proposing a 'Compute-Anchored Wage' bound where human wages are determined by compute capital costs rather than labor supply elasticity.