Google updates Android Bench with new LLMs, but Gemini still lags behind

Ars Technica Tools

Summary

Google updates Android Bench, a benchmark for LLMs in Android development tasks, adding new models like Claude Fable 5 and Qwen 3.7 Max, but its own Gemini models still trail behind competitors.

<p>Code generation is emerging as one of the most popular applications for large language models (LLMs), but not all agents are equally good at all development tasks. Google created a benchmark earlier this year to evaluate how LLMs perform in Android app development, and <a href="https://developer.android.com/bench">Android Bench</a> is getting a big update today. The leaderboard now includes a raft of new models, and Google has adopted a new framework that should be easier to use. Developers are invited to run their own tests and submit feedback that could shape the future of Android Bench.</p> <p>While they are popular coding tools, LLMs don't get everything right. Separating the useful outputs from straight-up slop means choosing the right tool. Android Bench aims to demonstrate which AI agents do best on a suite of 100 Android development tasks. After launching Android Bench in March, Google has added metrics like cost and efficiency, as well as open-weight models.</p> <p>To keep Android Bench relevant, Google is <a href="https://android-developers.googleblog.com/2026/07/android-bench-llm-measurement.html">updating the test</a> with eight new models, including all the latest heavy-hitters: Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus, and Qwen 3.7 Max.</p><p><a href="https://arstechnica.com/google/2026/07/google-revamps-android-ai-dev-benchmark-adds-fable-5-and-other-agents/">Read full article</a></p> <p><a href="https://arstechnica.com/google/2026/07/google-revamps-android-ai-dev-benchmark-adds-fable-5-and-other-agents/#comments">Comments</a></p>
Original Article
View Cached Full Text

Cached at: 07/09/26, 07:47 AM

# Google updates Android Bench with new LLMs, but Gemini still lags behind Source: [https://arstechnica.com/google/2026/07/google-revamps-android-ai-dev-benchmark-adds-fable-5-and-other-agents/](https://arstechnica.com/google/2026/07/google-revamps-android-ai-dev-benchmark-adds-fable-5-and-other-agents/) Code generation is emerging as one of the most popular applications for large language models \(LLMs\), but not all agents are equally good at all development tasks\. Google created a benchmark earlier this year to evaluate how LLMs perform in Android app development, and[Android Bench](https://developer.android.com/bench)is getting a big update today\. The leaderboard now includes a raft of new models, and Google has adopted a new framework that should be easier to use\. Developers are invited to run their own tests and submit feedback that could shape the future of Android Bench\. While they are popular coding tools, LLMs don’t get everything right\. Separating the useful outputs from straight\-up slop means choosing the right tool\. Android Bench aims to demonstrate which AI agents do best on a suite of 100 Android development tasks\. After launching Android Bench in March, Google has added metrics like cost and efficiency, as well as open\-weight models\. To keep Android Bench relevant, Google is[updating the test](https://android-developers.googleblog.com/2026/07/android-bench-llm-measurement.html)with eight new models, including all the latest heavy\-hitters: Claude Fable 5, Claude Sonnet 5, Claude Opus 4\.8, GLM 5\.2, Kimi K2\.7 Code, MiniMax M3, Qwen 3\.7 Plus, and Qwen 3\.7 Max\. Even the initial release of Android Bench didn’t have Google’s AI models at the top—OpenAI’s latest LLMs were slightly in the lead\. The story is worse for Gemini now that Google has expanded the lineup\. In the new leaderboard, Gemini 3\.1 Pro is in fifth place, behind GPT 5\.4, Claude Sonnet 5, and Claude Fable 5\. In fact, Fable 5 lives up to the hype with a sizeable lead at 84\.5 percent accuracy in the test\.

Similar Articles

Gemini 2.5: Our most intelligent models are getting even better

Google DeepMind Blog

Google announces Gemini 2.5 series updates, including improved 2.5 Pro and Flash models with new capabilities like Deep Think (enhanced reasoning mode), native audio output, and computer use abilities via Project Mariner. The models now lead on WebDev Arena and LMArena leaderboards.

Gemini 3.5 Flash Benchmarks

Reddit r/singularity

Benchmark results for the Gemini 3.5 Flash model are discussed, likely showcasing its performance across various AI tasks.

Fable 5 below even Gemini 3.1 on Livebench

Reddit r/singularity

A discussion on LiveBench results showing Fable 5 performing below Gemini 3.1, questioning whether the benchmark is flawed or Anthropic is optimizing for benchmarks.

Google updates its Gemini app to take on ChatGPT and Claude

TechCrunch AI

Google announces major updates to its Gemini app at Google I/O, including a Daily Brief feature, redesigned Neural Expressive interface, a personal AI agent called Gemini Spark, and integration of the new Gemini Omni video model to compete with ChatGPT and Claude.