Rippling's 2,100 scored runs experiment vs. the Stripe OpenRouter $7B deal

Reddit r/AI_Agents News

Summary

Rippling conducted a benchmark test of 15 AI models on payroll tasks, finding Anthropic's Opus 4.6 performed best but with a 9% failure rate, while Stripe acquired OpenRouter for $7B to help developers choose AI models.

Rippling ran a benchmark that included testing 15 AI models on real payroll work with 2,100 scored runs per model and pass/fail grading with unfinished = fail. Results: - 7 untuned models came in at 88.5%-89.5% - Z.ai's GLM 5.2 (open-source): 88.7% for $621 total - Anthropic's Opus 4.6 (only prompt-tuned model): 91.0% for $1,453 This was structured payroll work, i.e. API calls with strict rules and validation, and the winner Opus 4.6 failed 9% of the scored runs on this harness. Rippling built pass/fail grading and their own spend console around it. Weeks later, Stripe acquired OpenRouter for $7B. OpenRouter helps developers pick between AI models based on cost, latency, provider rate limits, region/compliance, or fit for the task. They process ~100 trillion tokens a month across 8M developers and earn ~5% commission on inference spend, about $140M ARR right now. The question is how much of actual agent inference is heavy on thinking and reasoning vs. rather simple structured output work? Or more directly, do we need smart routing work in the future, or is a simple role-based fixed setup sufficient?
Original Article

Similar Articles

Stripe Eyes $10 Billion Deal for AI Model Marketplace OpenRouter

Reddit r/LocalLLaMA

Stripe is in talks to acquire AI model marketplace OpenRouter for roughly $10 billion, marking a major expansion beyond payments. OpenRouter allows developers to access and switch between multiple AI models, and the deal would follow Stripe's broader push into AI infrastructure and its separate bid for PayPal.