@rohanpaul_ai: So much possibilities for on-device small models. Here @adrgrondin is running Google’s Gemma 4 E2B on iPhone 17 Pro. ~4…
Summary
Google's Gemma 4 E2B is demonstrated running on an iPhone 17 Pro via MLX optimization, achieving ~40 tokens/second with 128K context and offline thinking mode for coding and math.
View Cached Full Text
Cached at: 05/18/26, 08:29 AM
So much possibilities for on-device small models.
Here @adrgrondin is running Google’s Gemma 4 E2B on iPhone 17 Pro. ~40tk/s with MLX optimized for Apple Silicon SOTA coding & math on mobile with 128K context. Fully offline with thinking mode. https://t.co/sPK65qc7Bd
Similar Articles
@rohanpaul_ai: Gemma 4 (specifically its edge-optimized E2B and E4B variants) running fully offline on an iPhone via apps like Locally…
Google’s Gemma 4 E2B/E4B quantized variants now run fully offline on iPhone via apps like Locally AI, leveraging the Apple Neural Engine for on-device inference.
Gemma 4 26B A4B running on iPhone 17 Pro via model paging
Google's Gemma 4 26B A4B model can run on iPhone 17 Pro using model paging, enabling powerful on-device AI.
@googlegemma: Check out original post by @ivanfioravanti here:
A demonstration of the Gemma 4 E4B AI model running locally on an iPad using Apple MLX, showcased as an engaging application for children.
Introducing Gemma 3 270M: The compact model for hyper-efficient AI
Google introduces Gemma 3 270M, a compact 270-million parameter model designed for efficient on-device AI with strong instruction-following capabilities and extreme energy efficiency (0.75% battery for 25 conversations on Pixel 9 Pro).
@h100envy: Google engineer explained how to fine-tune a tiny LLM from 46% to 90% accuracy on your phone in 21 minutes - better tha…
A Google engineer shares a method to fine-tune a Gemma 270M model from 46% to 90% accuracy in 21 minutes on a phone, using synthetic data, LoRA, int4 quantization, achieving 2000 tokens per second offline.