on-device-ai

Tag

Cards List
#on-device-ai

Google AI Edge Gallery v1.0.13 & v1.0.14 updates: Gemma 4 Multi-Token Prediction, Pixel TPU support, experimental MCP, new skills, now saves chat history

Reddit r/LocalLLaMA · 2026-05-19 Cached

Google AI Edge Gallery v1.0.13 & v1.0.14 updates add support for Gemma 4 with multi-token prediction, Pixel TPU optimization, experimental MCP, new skills, and chat history saving, enhancing on-device generative AI capabilities.

0 favorites 0 likes
#on-device-ai

@ycombinator: General Instinct (@gen_instinct) deploys frontier AI models onto constrained edge hardware, helping robotics and physic…

X AI KOLs Following · 2026-05-19

General Instinct launches a deployment layer that enables frontier AI models to run on constrained edge hardware like Jetsons and mobile NPUs, helping robotics and physical AI teams achieve low-latency offline inference.

0 favorites 0 likes
#on-device-ai

@paulabartabajo_: The next AI boom won't be bigger data centers. It'll be compact intelligence running on the edge. You (and the planet) …

X AI KOLs Timeline · 2026-05-19 Cached

A tweet argues the next AI boom will be compact intelligence on edge devices rather than larger data centers, with Liquid AI supporting the vision of running AI on phones, cars, and everyday devices.

0 favorites 0 likes
#on-device-ai

RAG on Snapdragon X2 Laptop, 200K documents.

Reddit r/LocalLLaMA · 2026-05-15

VecML demonstrates its AI-PC software running RAG on 200K documents using the new Snapdragon X2 laptop, achieving low-token and low-memory retrieval. The software integrates multiple database functions into one platform, and controlled testing for macOS is now open.

0 favorites 0 likes
#on-device-ai

Gemma 4 + LiteRT-LM on mobile: much better memory/perf than my llama.cpp setup

Reddit r/LocalLLaMA · 2026-05-15

A user shares a hands-on comparison of running Gemma 4 with LiteRT-LM on mobile devices versus their previous llama.cpp setup, noting significantly better memory usage (1.5-2 GB vs 4-5 GB) and faster inference (2-4 seconds vs 7-10 seconds) on smartphones like Samsung S25 Ultra and iPhone 13 Pro Max.

0 favorites 0 likes
#on-device-ai

@berryxia: Great news for Mac users! Apple's on-device model advantage is back! I also saw today that Jina natively supports MLX in its framework! Previously, the release rhythm for open-source embedding models was usually like this: Day 0: Release PyTorch original. Day 7-30: Community converts to GGUF. Day 3…

X AI KOLs Timeline · 2026-05-13

Jina releases MLX-native embedding models simultaneously with PyTorch versions, highlighting the growing importance of Apple's MLX framework for local AI deployment.

0 favorites 0 likes
#on-device-ai

Needle: We Distilled Gemini Tool Calling Into a 26M Model

Reddit r/LocalLLaMA · 2026-05-12

Cactus-Compute released Needle, a 26M parameter open-source model distilled from Gemini for efficient on-device function calling using a novel Simple Attention Network architecture without MLPs.

0 favorites 0 likes
#on-device-ai

@JafarNajafov: Supertonic just killed ElevenLabs. A text-to-speech model that runs entirely on your device. No cloud. No API key. No p…

X AI KOLs Timeline · 2026-05-12

The article highlights Supertonic, an open-source text-to-speech model that runs entirely on-device, claiming superior speed and formatting accuracy compared to cloud-based services like ElevenLabs and OpenAI.

0 favorites 0 likes
#on-device-ai

@Prince_Canuma: My @aiDotEngineer talk is live: "On-device Intelligence using MLX" Huge thanks to @swyx and the team for having me — ha…

X AI KOLs Following · 2026-05-11

The author announces their live talk titled 'On-device Intelligence using MLX' at the aiDotEngineer event, expressing gratitude to the organizers and community contributors.

0 favorites 0 likes
#on-device-ai

Show HN: TikTok but for Scientific Papers

Hacker News Top · 2026-05-11 Cached

Papel is a new research-focused social platform that leverages AI-powered vector search and on-device RAG to help researchers discover, discuss, and quiz themselves on academic papers. It offers personalized feeds, local AI chat via Apple Intelligence or MLX, and gamified learning features.

0 favorites 0 likes
#on-device-ai

Built a practical voice-first AI tool for ADHD/executive dysfunction — one-tap brain dump → structured reminders & tasks (not a full autonomous agent)

Reddit r/AI_Agents · 2026-05-10

The author introduces SAVI, an iOS app designed for ADHD users that converts voice brain dumps into structured tasks and reminders using on-device AI like Whisper and GPT-4o.

0 favorites 0 likes
#on-device-ai

Chrome’s AI features may be hogging 4GB of your computer storage

Lobsters Hottest · 2026-05-09 Cached

Google Chrome is automatically downloading a 4GB Gemini Nano model weights file to users' devices to power on-device AI features like scam detection and writing assistance, often without clear notification about storage requirements. Users can disable the On-Device AI toggle in Chrome settings to remove the file and prevent re-downloads.

0 favorites 0 likes
#on-device-ai

@garrytan: Downloading now... 1M token context window with supposedly usable coding agent capability all on a 128GB Macbook Pro is

X AI KOLs Following · 2026-05-09 Cached

Garry Tan highlights a model with a 1M token context window and coding agent capabilities running locally on a 128GB MacBook Pro, expressing excitement about the milestone.

0 favorites 1 likes
#on-device-ai

@rohanpaul_ai: atomic[.]chat just made Gemma 4 26B faster inside LLaMA.cpp. making token generation about 40% faster in its MacBook Pr…

X AI KOLs Following · 2026-05-07

atomic.chat has optimized Gemma 4 26B inference in LLaMA.cpp, achieving ~40% faster token generation on MacBook Pro M5 Max using Multi-Token Prediction (MTP) speculative decoding. This is a notable win for local AI users running desktop apps, coding agents, and private on-device assistants.

0 favorites 0 likes
#on-device-ai

X-OmniClaw Technical Report: A Unified Mobile Agent for Multimodal Understanding and Interaction

Hugging Face Daily Papers · 2026-05-07 Cached

This technical report introduces X-OmniClaw, a unified mobile agent system designed for multimodal understanding and interaction on Android devices. It details the architecture for perception, memory management, and action execution using on-device AI capabilities.

0 favorites 0 likes
#on-device-ai

Supertone/supertonic-3

Hugging Face Models Trending · 2026-05-06 Cached

Supertonic 3 is a lightweight, open-weight text-to-speech model designed for fast on-device inference, expanding support to 31 languages with improved stability and expression tags.

0 favorites 0 likes
#on-device-ai

Enabling privacy-preserving AI training on everyday devices

MIT News — Artificial Intelligence · 2026-04-29 Cached

MIT researchers developed a new framework called FTTE that accelerates privacy-preserving federated learning by 81%, enabling efficient AI training on resource-constrained edge devices like smartwatches and sensors.

0 favorites 0 likes
#on-device-ai

AngelSlim/Hy-MT1.5-1.8B-1.25bit

Hugging Face Models Trending · 2026-04-28 Cached

Tencent's AngelSlim team released Hy-MT1.5-1.8B-1.25bit, a highly compressed 1.25-bit machine translation model supporting 33 languages that fits in 440MB for on-device use. It utilizes the Sherry quantization algorithm to achieve world-class translation quality comparable to much larger models.

1 favorites 1 likes
#on-device-ai

google/gemma-4-31B-it-assistant

Hugging Face Models Trending · 2026-04-23 Cached

Google DeepMind releases Gemma 4, a family of open-weights multimodal models featuring Multi-Token Prediction (MTP) for up to 2x decoding speedups, supporting text, image, video, and audio with enhanced reasoning and coding capabilities.

0 favorites 0 likes
#on-device-ai

What impedes apps using AI to make the user’s device the server running a local LLM?

Reddit r/singularity · 2026-04-22

A user reflects on why more apps don’t run local LLMs directly on phones, noting Gemma 2-4B models already work offline and could eliminate server costs while maintaining near-GPT-4o quality.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback