Tag
HybridInfer introduces a thermal-aware reinforcement-learning router for multi-tier LLM inference, achieving higher quality and cost efficiency than heuristics on real Android devices.
Qwen Intelligence launches three state-of-the-art mobile AI agents: Mobile Planner Agent for task decomposition, Mobile-Use Agent with high end-to-end success rates, and Mobile Creative Agent for rapid content generation, alongside an open benchmark suite for evaluation.
TierKV proposes a predictive multi-tier KV caching framework to optimize memory usage and throughput for long-context LLMs on mobile devices, achieving significant performance improvements with minimal accuracy degradation.
Edge0-35B-A3B-preview is a sparse MoE model that enables efficient AI inference on mobile devices by using streaming expert offloading and quantization, achieving 15 tok/s with under 3 GiB of memory.
SnapBench introduces a paired corruption benchmark for robust snap-and-ask multimodal retrieval on mobile, revealing image noise as a key degradation factor and proposing an adaptive fusion method for modality reliability.
A tweet speculates that Apple, Qualcomm, and Google could become leading AI companies by focusing on portable, edge AI devices, referencing Google Gemma running on NVIDIA Jetson.
The article describes the experience of running an AI agent on an Ubuntu Touch smartphone, which is always on with access to sensors, allowing continuous conversation via Telegram.
Google DeepMind announced SL2T, a system developed with the Deaf community to bring ASL input to phones, with plans to expand to more sign languages and applications.
Creator of bigedgeonmoe open-source codebase enables running massive MoE models (up to 120B parameters) on mobile devices and consumer PCs, achieving 6 tokens/s for Qwen 35B on a mid-range phone.
This paper empirically studies how prompt wording affects energy consumption for on-device LLMs, showing that keyword choices can significantly impact decoding length and total energy, suggesting prompt engineering as a lightweight energy optimization lever.
LightMem-Ego is a lightweight streaming multimodal memory system for everyday life assistance that continuously captures egocentric visual and audio streams, organizes them into hierarchical memory, and retrieves grounded answers to user queries about past experiences, deployable on smartphones and AI glasses.
PrismML, a Khosla-backed startup, releases a compressed version of Qwen-3.6-27B, claiming it's the largest AI model ever to run on an iPhone.
The PyTorch Foundation supported the ExecuTorch Hackathon in San Francisco, where over 100 participants built real-time on-device AI applications using PyTorch and ExecuTorch on Snapdragon-powered Samsung Galaxy S25 Ultra devices. Winning projects included SafeScreen AI, SixthSense, and Toddle AI, showcasing local execution benefits for responsiveness, privacy, and offline capability.
Announces Acti, an agentic keyboard that turns the passive keyboard into an active interface capable of executing actions directly within apps, such as pulling Notion docs or generating mini apps.
Running Gemma 12B model on a Google Pixel 10 Pro using llama.cpp achieves 6.5 tokens per second prompt processing and 1.3 tokens per second generation with under 10 watts power consumption, demonstrating efficient on-device AI inference.
This paper presents the first end-to-end RAG pipeline running entirely on a mobile NPU (Qualcomm Hexagon on Snapdragon X Elite), achieving up to 18x faster LLM prefilling and 4x lower energy vs. CPU, with no quality regression.
Benchmark shows local Stable Diffusion 1.5 on iPhone can generate 512x512 images in as little as 3.1 seconds using optimized models like Realistic Vision V5.1 Hyper, making on-device AI image generation practical.
This article discusses the imminent arrival of AI-powered smartphones and the implications for consumers and the tech industry.
This article argues that the real issue with integrating Gemini deeper into Android isn't just privacy, but the action boundary—what the AI can read, suggest, draft, change, send, buy, or delete—and proposes a tiered consent model for different levels of AI agency.
Google and Apple are bringing AI-powered 'vibe coding' to mobile, allowing users to create custom Android apps, widgets, and shortcuts via natural language prompts, as demonstrated at Google I/O 2026 and reported for iOS.