Tag
Meta introduced Muse Glimmer, an open-weight 30B-parameter model distilled from Muse Spark for on-device agentic workflows, with ExecuTorch now supporting running it on NVIDIA GPUs and Apple silicon.
Meta announces a new open-source model optimized for on-device deployment, aiming to bring efficient AI inference to edge devices.
Introduces GRASP, a method that uses Group Relative Policy Optimization to train a small on-device language model for adversarial anonymization, improving the privacy-utility trade-off over DPO-distilled baselines while running at a fraction of the cost of frontier teacher models.
The author shares hands-on experience with LFM 2.6B, a small model designed for phones, praising its speed and usefulness for quick tasks like summarization and autocomplete, though it has a 128k context limit.
NVIDIA's entire speech stack—ASR, TTS, and codec—is now quantized to GGUF and runs locally on-device via NeMo-Speech.cpp, with new model releases for Magpie-TTS, Nemotron Speech Streaming, and Parakeet.
Maxime Labonne highlights early fine-tunes of LFM2.5-2.6B, while Bad Theory Labs releases two open-weight models: BTL-4 35B, a frontier agentic reasoning model, and Macaw 2.7B, an on-device Mac agent, with notable BFCL v4 performance.
VisionPsy-Nano-460M-Flash is a new 460M vision-language model that uses only 64 visual tokens per image, cutting first-token latency to 0.3s on iPhone while retaining roughly 99% of the full model's benchmark score. The optimization trades off some OCR and fine-detail quality.
MemArena is a new ego-centric benchmark for evaluating on-device personal memory assistants, using a MASim agent simulator to generate multi-session conversational worlds and ground truth across recall, reasoning, and trustworthiness dimensions. Initial results show memory-backend choice often matters more than reader scale, and permission-aware access remains a universal challenge.
VibeVoice 1.5B runs locally on an iPhone with ~2.2 GB memory and up to 1.28× real-time speed; the author plans to release the xcframework and code for audio.cpp.
Liquid AI releases LFM2.5-2.6B, a compact agentic model designed to run entirely on-device, enabling free inference, low latency, and privacy. The post details its training pipeline including SFT, teacher specialization, distillation, and agentic RL.
DeepGrove releases Maple-Preview, an open-source 20B-A1B ternary-weight reasoning LLM with SOTA reasoning for its weight class, capable of 200+ tokens/sec on a Mac mini M4 and competitive with larger models.
Maple-Preview is a ternary 20B MoE model that runs at 120 tokens per second on an iPhone, showcasing efficient on-device inference.
NewEyes is a consumer AI app from Collov Labs that uses on-device multimodal intelligence to make the camera understand and act on the world—enabling visual shopping, style, room design, calorie counting, and more.
Liquid AI released LFM2.5-2.6B, an on-device agentic model that plans, calls tools, and works through multi-step tasks on phones, laptops, PCs, and robots, with data never leaving the device.
Liquid AI releases LFM2.5-2.6B, a compact agentic model designed for on-device deployment, supporting tool calling and multi-step workflows with efficient inference on CPUs and GPUs.
Open Minis is an on-device AI agent that runs on your phone, designed to be open and secure.
Zen Whisper is a Mac app offering on-device dictation that can type into any application.
This paper proposes CoRA, a gradient-free framework for task-conditioned retrieval in on-device in-context learning, using frozen encoders and closed-form ridge regression to build compact retrieval bases without fine-tuning or backpropagation.
Liquid AI's head of post-training explains how to build a sub-1GB on-device model in 20 minutes using LFM2.5, on-policy preference alignment, agentic RL, curriculum training, and iterative model merging, achieving tool-calling reliability that beats much larger models.
A free on-device application that lets you talk to coding or AI agents with voice responses, take screenshots, and connect meeting recordings without API keys.