on-device

Tag

Cards List
#on-device

@PyTorch: Today @AIatMeta introduced Muse Glimmer, an open-weight, 30-billion-parameter model distilled from Meta’s Muse Spark fo…

X AI KOLs Following · 5h ago Cached

Meta introduced Muse Glimmer, an open-weight 30B-parameter model distilled from Muse Spark for on-device agentic workflows, with ExecuTorch now supporting running it on NVIDIA GPUs and Apple silicon.

0 favorites 0 likes
#on-device

Meta releases new on-device optimized open source model

Reddit r/singularity · 9h ago

Meta announces a new open-source model optimized for on-device deployment, aiming to bring efficient AI inference to edge devices.

0 favorites 0 likes
#on-device

GRASP: Reinforcing Language Model Anonymizers with Group Relative Policy Optimization

arXiv cs.CL · 15h ago Cached

Introduces GRASP, a method that uses Group Relative Policy Optimization to train a small on-device language model for adversarial anonymization, improving the privacy-utility trade-off over DPO-distilled baselines while running at a fraction of the cost of frontier teacher models.

0 favorites 0 likes
#on-device

LFM 2.6B is a lot of fun.

Reddit r/LocalLLaMA · yesterday

The author shares hands-on experience with LFM 2.6B, a small model designed for phones, praising its speed and usefulness for quick tasks like summarization and autocomplete, though it has a 128k context limit.

0 favorites 0 likes
#on-device

🟩 NVIDIA's whole speech stack just went local. ASR + TTS + codec, quantized to GGUF, running on-device via NeMo-Speech.cpp

Reddit r/LocalLLaMA · 3d ago

NVIDIA's entire speech stack—ASR, TTS, and codec—is now quantized to GGUF and runs locally on-device via NeMo-Speech.cpp, with new model releases for Magpie-TTS, Nemotron Speech Streaming, and Parakeet.

0 favorites 0 likes
#on-device

@maximelabonne: It didn't take two days to get the first fine-tunes of LFM2.5-2.6B. This one looks cool!

X AI KOLs Following · 4d ago Cached

Maxime Labonne highlights early fine-tunes of LFM2.5-2.6B, while Bad Theory Labs releases two open-weight models: BTL-4 35B, a frontier agentic reasoning model, and Macaw 2.7B, an on-device Mac agent, with notable BFCL v4 performance.

0 favorites 0 likes
#on-device

A 460M VLM gets first-token latency down to 0.3s on an iPhone by using only 64 visual tokens

Reddit r/LocalLLaMA · 5d ago

VisionPsy-Nano-460M-Flash is a new 460M vision-language model that uses only 64 visual tokens per image, cutting first-token latency to 0.3s on iPhone while retaining roughly 99% of the full model's benchmark score. The optimization trades off some OCR and fine-detail quality.

0 favorites 0 likes
#on-device

MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale

arXiv cs.CL · 5d ago Cached

MemArena is a new ego-centric benchmark for evaluating on-device personal memory assistants, using a MASim agent simulator to generate multi-session conversational worlds and ground truth across recall, reasoning, and trustworthiness dimensions. Initial results show memory-backend choice often matters more than reader scale, and permission-aware access remains a universal challenge.

0 favorites 0 likes
#on-device

VibeVoice 1.5B Running Locally...On an iPhone! Only ~2.2 GB of Memory and Up to 1.28× Real-Time Speed

Reddit r/LocalLLaMA · 5d ago

VibeVoice 1.5B runs locally on an iPhone with ~2.2 GB memory and up to 1.28× real-time speed; the author plans to release the xcframework and code for audio.cpp.

0 favorites 0 likes
#on-device

LFM2.5-2.6B: Deploy Agents Everywhere (8 minute read)

TLDR AI · 5d ago Cached

Liquid AI releases LFM2.5-2.6B, a compact agentic model designed to run entirely on-device, enabling free inference, low latency, and privacy. The post details its training pipeline including SFT, teacher specialization, distillation, and agentic RL.

0 favorites 0 likes
#on-device

deepgrove/maple-preview

Hugging Face Models Trending · 5d ago Cached

DeepGrove releases Maple-Preview, an open-source 20B-A1B ternary-weight reasoning LLM with SOTA reasoning for its weight class, capable of 200+ tokens/sec on a Mac mini M4 and competitive with larger models.

0 favorites 0 likes
#on-device

Show HN: Maple-Preview – ternary 20B MoE running at 120 tok/s on a iPhone

Hacker News Top · 5d ago

Maple-Preview is a ternary 20B MoE model that runs at 120 tokens per second on an iPhone, showcasing efficient on-device inference.

0 favorites 0 likes
#on-device

@CollovLabs: What if your camera could… Shop what you see. Style what you wear. Design the room around you. Count the calories on yo…

X AI KOLs Following · 6d ago Cached

NewEyes is a consumer AI app from Collov Labs that uses on-device multimodal intelligence to make the camera understand and act on the world—enabling visual shopping, style, room design, calorie counting, and more.

0 favorites 0 likes
#on-device

@maximelabonne: LFM2.5-2.6B is available today on @huggingface First agentic model of its kind, it destroys our previous release on EVE…

X AI KOLs Following · 6d ago Cached

Liquid AI released LFM2.5-2.6B, an on-device agentic model that plans, calls tools, and works through multi-step tasks on phones, laptops, PCs, and robots, with data never leaving the device.

0 favorites 0 likes
#on-device

Deploy local agents everywhere with LFM2.5-2.6B

Hugging Face Blog · 6d ago Cached

Liquid AI releases LFM2.5-2.6B, a compact agentic model designed for on-device deployment, supporting tool calling and multi-step workflows with efficient inference on CPUs and GPUs.

0 favorites 0 likes
#on-device

Open Minis

Product Hunt · 2026-08-02

Open Minis is an on-device AI agent that runs on your phone, designed to be open and secure.

0 favorites 0 likes
#on-device

Zen Whisper

Product Hunt · 2026-07-31

Zen Whisper is a Mac app offering on-device dictation that can type into any application.

0 favorites 0 likes
#on-device

Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning

arXiv cs.CL · 2026-07-31 Cached

This paper proposes CoRA, a gradient-free framework for task-conditioned retrieval in on-device in-context learning, using frozen encoders and closed-form ridge regression to build compact retrieval bases without fine-tuning or backpropagation.

0 favorites 0 likes
#on-device

@h100envy: Liquid AI's head of post-training explained how they built a small model that runs on-device under 1 GB in 20 minutes -…

X AI KOLs Following · 2026-07-30 Cached

Liquid AI's head of post-training explains how to build a sub-1GB on-device model in 20 minutes using LFM2.5, on-policy preference alignment, agentic RL, curriculum training, and iterative model merging, achieving tool-calling reliability that beats much larger models.

0 favorites 0 likes
#on-device

Free on-device voice control for AI agents where agents can talk back as well

Reddit r/AI_Agents · 2026-07-30

A free on-device application that lets you talk to coding or AI agents with voice responses, take screenshots, and connect meeting recordings without API keys.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback