Un modello 100% locale, anche sul tuo smarphone!
Summary
Rilasciata un'interfaccia per gestire due piccoli modelli (4B e 1.7B) che girano localmente sullo smartphone. Il 4B funziona bene su telefoni di fascia alta; il 1.7B ha problemi di stabilità con il reasoning, in fase di miglioramento con fine-tuning approfondito usando 130k esempi e distillazione da un teacher 32B.
Similar Articles
Un modello linguistico locale, privato 100%, sul tuo smartphone!!
Un modello linguistico locale e privato (Qwen 3 da 1.5B e 4B quantizzati) può girare offline su smartphone, con fine-tuning e LoRA distillato da un 32B.
Un modello 100% locale sul tuo smartphone!
A developer shares their success in fine-tuning Qwen 3 models (1.5B and 4B) for local use on smartphones, with a downloadable APK that works offline, and plans for a Windows version.
Getting real work out of a 4B local model: the distill-on-idle pipeline behind an on-device "memory" assistant
Describes a 'distill-on-idle' pipeline that enables a 4B parameter local model to run effectively as an on-device memory assistant, demonstrating practical use of small models.
@rohanpaul_ai: So much possibilities for on-device small models. Here @adrgrondin is running Google’s Gemma 4 E2B on iPhone 17 Pro. ~4…
Google's Gemma 4 E2B is demonstrated running on an iPhone 17 Pro via MLX optimization, achieving ~40 tokens/second with 128K context and offline thinking mode for coding and math.
@AdinaYakup: MiniCPM V4.6 a 1B MLLM that actually runs on your phone, just released by @OpenBMB 1B - Apache2.0 Runs on iOS, Android,…
OpenBMB has released MiniCPM V4.6, a 1B-parameter multimodal large language model optimized for mobile devices under the Apache 2.0 license. It features mixed visual token compression and claims approximately 1.5x faster throughput than Qwen3.5 0.8B while running natively on iOS, Android, and HarmonyOS.