A solo developer tests Opus 5, Opus 4.8, GPT-5.6 Sol and Kimi K3 via a multi-model router with free credit, discovering that evaluation budgets and input preprocessing matter more than raw model choice.
Building solo, and my wall had nothing to do with my code: I'd become scared of my own experiments. Every pipeline change meant re-running the whole batch to know if I'd improved it or broken it. Opus 5 and Opus 4.8 are both $5 in / $25 out. Sol is cheaper but has a million token context I kept filling. So every honest test cost money, and I quietly stopped testing and started guessing. With nobody reviewing your diffs, that's the worst failure mode there is. Tried one of those multi-model routers expecting a bait and switch. Signup credit, no card, all three plus Kimi K3 behind one OpenAI-compatible URL. (Weight my enthusiasm accordingly: these run referral programs. No link, I get nothing from this.) Not a scam. But it didn't work how I expected, and that's the useful part. I thought free credit meant free compute. What it actually bought was an evaluation budget: one real batch, every candidate once, outputs side by side, pick one, commit, stop shopping. Three things I didn't see coming: Reasoning level moved my bill more than model choice ever did. Opus 5 thinks by default, right for a nasty bug, quietly expensive for a find-and-replace. My biggest win came from preprocessing the input before the model saw it, not a stronger model. Never would've found that while I was too scared to compare. Capability and instruction-following are separate axes. The strongest model isn't automatically the one you want in your repo when you're the only reviewer. I had frontier models write a gorgeous plan, list the files they were about to edit, then stop and bill me for the thinking. Model or router plumbing? Genuinely can't tell. Production critical, go direct. The real fix wasn't the money. It's that I measure things again. Happy to get into the setup or the eval batch, just keep it in the thread rather than DMs. How are you handling this: switching by hand, one router, or picked one and eating the cost?
Matt Shumer tests GPT-6 Sol and shares his preference for Astra/Fable 5.1 and Opus 5.5 models, while referencing OpenAI's announcement of faster and more affordable GPT-6 Sol and Luna models.
A tweet from Philip Kiely highlights cost savings by switching from closed-source AI models to open-source alternatives, using Baseten's ROI calculator tool.
Alexander Yue introduces a new browser-use benchmark where Opus 5 and GPT-5.6 Sol show similar performance, emphasizing the benchmark's robust design with verified rubrics for LLM judges.