mobile-agents

Tag

Cards List
#mobile-agents

The VMs Powering Mobile Agents (Instinct, Claude Code)

Hacker News Top · 6d ago Cached

The article details the virtual machine technologies powering mobile AI agents, such as Claude Code's Firecracker microVM and Instinct's use of E2B sandboxes, highlighting their security and lifecycle management.

0 favorites 0 likes
#mobile-agents

@omooretweets: Agents that actually work (esp. on mobile) are going to be a huge tailwind for voice AI I've massively increased attemp…

X AI KOLs Timeline · 2026-08-30 Cached

The tweet discusses how functional agents on mobile could drive voice AI growth, with the author noting increased voice dictate use due to Instinct, and references Jane Manchun Wong's tweet about Instinct's 'Talk to Instinct' feature for iPhone.

0 favorites 0 likes
#mobile-agents

I think we might have been thinking about mobile agents the wrong way. APIs give you control, but understanding the screen gives you something else.

Reddit r/AI_Agents · 2026-08-29

The author questions traditional mobile agent approaches relying on API access and proposes using screen understanding via hardware like aiden-firmware to enable more general-purpose agents that interact like humans.

0 favorites 0 likes
#mobile-agents

Automated Trajectory Evaluation for Mobile Agents via Step-Level Consequence Reasoning and Aggregation

arXiv cs.AI · 2026-08-24 Cached

The paper introduces CRATE, a two-stage framework using step-level consequence reasoning to evaluate mobile agents, achieving high F1-scores on benchmarks like AndroidWorld and MobileRisk.

0 favorites 0 likes
#mobile-agents

CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications

arXiv cs.AI · 2026-08-13 Cached

CoAdapt-GUI is a test-time adaptation framework for mobile GUI agents that jointly adapts workflow context and policy, improving performance on unseen-app benchmarks like AndroidWorld-Generalization and AndroidWorld Plus.

0 favorites 0 likes
#mobile-agents

Benchmarking LLM Judges for Mobile Agent Evaluation

arXiv cs.AI · 2026-08-13 Cached

This paper introduces MobileJudgeBench, a benchmark with 931 human-annotated trajectories for systematically evaluating LLM-based judges on mobile agent tasks. It finds that simple baseline judges with sampled screenshots rival purpose-built methods, with the LLM backbone being the primary driver of quality.

0 favorites 0 likes
#mobile-agents

AndroidReality: How Far Are Mobile Agents from the Real World?

arXiv cs.AI · 2026-08-11 Cached

Introduces AndroidReality, a perturbation-based framework for evaluating and improving the robustness of mobile agents, with a taxonomy of real-world interface perturbations and a training-free Test-Time Introspective Recovery (TTIR) mechanism.

0 favorites 0 likes
#mobile-agents

iOSWorld: A Benchmark for Personally Intelligent Phone Agents

Hugging Face Daily Papers · 2026-06-08 Cached

Introduces iOSWorld, an interactive native iOS simulator benchmark with persistent user identity across 26 apps, designed to evaluate personalized mobile agent capabilities through 133 tasks of increasing difficulty.

0 favorites 0 likes
#mobile-agents

MIRAGE: Mobile Agents with Implicit Reasoning and Generative World Models

arXiv cs.AI · 2026-06-04 Cached

MIRAGE is a framework for mobile GUI agents that replaces verbose chain-of-thought reasoning with compact continuous latent representations, incorporating a generative world model perspective to predict future screen states before acting. On AndroidWorld and AndroidControl benchmarks, it achieves competitive or superior performance while reducing generated tokens by over 75%.

0 favorites 0 likes
#mobile-agents

Perceive Before Reasoning: A Pre-Reasoning Perception Framework for Efficient and Reliable Proactive Mobile Agents

arXiv cs.AI · 2026-06-03 Cached

This paper proposes a Pre-Reasoning Perception Framework (PRPF) for proactive mobile agents, decoupling intervention timing from assistance generation to improve efficiency and reduce false triggers.

0 favorites 0 likes
#mobile-agents

Is state tracking the hardest part of phone-use AI?

Reddit r/AI_Agents · 2026-05-20

The author observes that the hardest part of phone-use AI agents is tracking state changes, as mobile interfaces have more dynamic and interruptive UI changes compared to desktop, and asks for others' experience.

0 favorites 0 likes
← Back to home

Submit Feedback