AndroidLife: Can an AI agent survive a day in the life of a real user? Qwen3.8-27b run: 56.7% SR

Reddit r/LocalLLaMA News

Summary

An AI agent using the Qwen3.8-27b model from Alibaba achieved a 56.7% success rate in the AndroidLife benchmark, running 60 real tasks on a OnePlus phone with detailed performance and efficiency metrics.

I let AI run my phone 60 real tasks, back to back, on the OnePlus I use every day - Best text model still failed 43% of them - Peak chip temp 98.2 C - 69% of the battery gone The benchmark is AndroidLife, and this is the first of 11 models, qwen3.8-27b from Alibaba, running in text mode Success 56.7%, so it failed 43% of the 60 tasks 29.25 steps and about 6 minutes per task $0.118 per task 69% of the battery gone Peak chip 98.2 C, power-amp/skin 48.9 C, battery 37.9 C Asked twice on the 11 ASK USER tasks Best text score on the board, and still under 6 in 10 Where it loses points Buckets: easy 80.8%, medium 52.9%, hard 23.5% Single apps it mostly handles, but three apps in a row and it taps in circles until the step limit runs out, which is where most of the 43% goes Three claimed wins did not hold up on the device, a calendar called clash free while two events overlapped, travel times it never opened, and a name it picked instead of asking On the two sets of planted tasks 11 ASK USER tasks per run, split 7 SINGLE and 4 MULTI, they hide a fact only I know so the agent has to ask me before it can finish SINGLE is one missing fact, MULTI needs several questions to clear up ambiguity, and it asked twice, never on the MULTI ones 7 hallucination controls per run, tasks where the data is deliberately missing, and the only right answer is to say the thing is not there It handled 4 of the 7 honestly and invented nothing What AndroidLife is 60 public tasks, from a corpus of 530 across 28 days and 31 apps Real phone over wifi, no emulator Graded on device state, never the model's own report Seeded accounts and placeholder data, no banking apps The best text run right now is 56.7%, $0.118 a task, 6 minutes a task, and a phone that finished at 98.2 C with 69% of its charge gone Full leaderboard: https://androidlife-website.vercel.app/ All 60 public tasks and this model's trajectories: https://androidlife-website.vercel.app/pages/tasks-public.html?run=qwen-28
Original Article

Similar Articles

Qwen3.7: The Agent Frontier (15 minute read)

TLDR AI

Alibaba's Qwen team has released Qwen3.7-Max, a proprietary agent-foundation model achieving top scores on multiple benchmarks including Terminal-Bench 2.0, SWE-Pro, and GPQA Diamond, with consistent performance across various code environments.