Google updates Android Bench, a benchmark for LLMs in Android development tasks, adding new models like Claude Fable 5 and Qwen 3.7 Max, but its own Gemini models still trail behind competitors.
<p>Code generation is emerging as one of the most popular applications for large language models (LLMs), but not all agents are equally good at all development tasks. Google created a benchmark earlier this year to evaluate how LLMs perform in Android app development, and <a href="https://developer.android.com/bench">Android Bench</a> is getting a big update today. The leaderboard now includes a raft of new models, and Google has adopted a new framework that should be easier to use. Developers are invited to run their own tests and submit feedback that could shape the future of Android Bench.</p>
<p>While they are popular coding tools, LLMs don't get everything right. Separating the useful outputs from straight-up slop means choosing the right tool. Android Bench aims to demonstrate which AI agents do best on a suite of 100 Android development tasks. After launching Android Bench in March, Google has added metrics like cost and efficiency, as well as open-weight models.</p>
<p>To keep Android Bench relevant, Google is <a href="https://android-developers.googleblog.com/2026/07/android-bench-llm-measurement.html">updating the test</a> with eight new models, including all the latest heavy-hitters: Claude Fable 5, Claude Sonnet 5, Claude Opus 4.8, GLM 5.2, Kimi K2.7 Code, MiniMax M3, Qwen 3.7 Plus, and Qwen 3.7 Max.</p><p><a href="https://arstechnica.com/google/2026/07/google-revamps-android-ai-dev-benchmark-adds-fable-5-and-other-agents/">Read full article</a></p>
<p><a href="https://arstechnica.com/google/2026/07/google-revamps-android-ai-dev-benchmark-adds-fable-5-and-other-agents/#comments">Comments</a></p>
# Google updates Android Bench with new LLMs, but Gemini still lags behind
Source: [https://arstechnica.com/google/2026/07/google-revamps-android-ai-dev-benchmark-adds-fable-5-and-other-agents/](https://arstechnica.com/google/2026/07/google-revamps-android-ai-dev-benchmark-adds-fable-5-and-other-agents/)
Code generation is emerging as one of the most popular applications for large language models \(LLMs\), but not all agents are equally good at all development tasks\. Google created a benchmark earlier this year to evaluate how LLMs perform in Android app development, and[Android Bench](https://developer.android.com/bench)is getting a big update today\. The leaderboard now includes a raft of new models, and Google has adopted a new framework that should be easier to use\. Developers are invited to run their own tests and submit feedback that could shape the future of Android Bench\.
While they are popular coding tools, LLMs don’t get everything right\. Separating the useful outputs from straight\-up slop means choosing the right tool\. Android Bench aims to demonstrate which AI agents do best on a suite of 100 Android development tasks\. After launching Android Bench in March, Google has added metrics like cost and efficiency, as well as open\-weight models\.
To keep Android Bench relevant, Google is[updating the test](https://android-developers.googleblog.com/2026/07/android-bench-llm-measurement.html)with eight new models, including all the latest heavy\-hitters: Claude Fable 5, Claude Sonnet 5, Claude Opus 4\.8, GLM 5\.2, Kimi K2\.7 Code, MiniMax M3, Qwen 3\.7 Plus, and Qwen 3\.7 Max\.
Even the initial release of Android Bench didn’t have Google’s AI models at the top—OpenAI’s latest LLMs were slightly in the lead\. The story is worse for Gemini now that Google has expanded the lineup\. In the new leaderboard, Gemini 3\.1 Pro is in fifth place, behind GPT 5\.4, Claude Sonnet 5, and Claude Fable 5\. In fact, Fable 5 lives up to the hype with a sizeable lead at 84\.5 percent accuracy in the test\.
Google announces Gemini 2.5 series updates, including improved 2.5 Pro and Flash models with new capabilities like Deep Think (enhanced reasoning mode), native audio output, and computer use abilities via Project Mariner. The models now lead on WebDev Arena and LMArena leaderboards.
Google may be testing an upgraded Gemini Flash model on LM Arena, showing incremental improvements over the current version, with possible naming as Gemini 3.6 Flash or Gemini 4 Flash.
A discussion on LiveBench results showing Fable 5 performing below Gemini 3.1, questioning whether the benchmark is flawed or Anthropic is optimizing for benchmarks.
Google announces major updates to its Gemini app at Google I/O, including a Daily Brief feature, redesigned Neural Expressive interface, a personal AI agent called Gemini Spark, and integration of the new Gemini Omni video model to compete with ChatGPT and Claude.