Ranked AI models by what people actually use instead of benchmark scores - the benchmark champion barely makes the top 20
Summary
A ranking of AI models by real usage, cost, and speed reveals that benchmark champions often trail in actual adoption, with cheaper/faster models like Flash Lite and GPT-5 leading over premium counterparts like Gemini 3.1 Pro.
Similar Articles
The "One-Size-Fits-All" AI era is dead. I benchmarked GPT-5.5, Claude 4.7, Gemini 3.1 Pro, and DeepSeek V4 Pro here is the actual state of the frontier.
A benchmarking analysis of GPT-5.5, Claude Opus 4.7, Gemini 3.1 Pro, and DeepSeek V4 Pro reveals that no single model dominates all tasks; optimal performance requires a multi-model router with specialized model usage based on strengths and weaknesses.
Some models got priced the same for a week, so I watched what people actually used
When several AI models were priced equally for a week, actual token usage revealed preference differences from leaderboard rankings, showing that coding and general chat have different top models and long context usage concentrated on two trusted models.
GPT-5, the world best model just 1 year ago, is today inferior to Qwen3.6 27B and most today’s low-tier models
GPT-5, which was the best model just a year ago, is now outperformed by Qwen3.6 27B and other current low-tier models, highlighting the rapid pace of AI advancement.
The best model is the one you can actually run
The article argues that the best AI model is not necessarily the most powerful, but the one that can be practically deployed and run efficiently, emphasizing the importance of considering real-world constraints like cost and hardware requirements.
SemiAnalysis: Gemini 3.8 Flash and Muse Spark 1.3 are two of the most clearly benchmaxxed models we've seen yet.
SemiAnalysis reports that Gemini 3.8 Flash and Muse Spark 1.3 are among the most clearly benchmarked AI models, showcasing strong performance.