on-device-inference

Tag

Cards List
#on-device-inference

Ling-3.0-flash MXFP4 released and running locally on one DGX Spark.

Reddit r/LocalLLaMA · 3d ago

Ling-3.0-flash MXFP4, a quantized model, has been released and runs locally on a single DGX Spark, achieving ~80 tok/s decoding and 2,500-3,500 tok/s long-input prefilling, enabling private on-device inference for coding, agents, and offline batch jobs.

0 favorites 0 likes
#on-device-inference

MacPaw taps Liquid AI to offer on-device inference to devs building for its app store

TechCrunch AI · 3d ago Cached

MacPaw partners with Liquid AI to bring on-device AI inference and local memory to its products and app store, planning to offer the tech stack to developers and introduce credit-based AI pricing.

0 favorites 0 likes
#on-device-inference

@googledevs: Meet LiteRT.js: @Google’s new Edge AI runtime for the web! We've made it easier to convert from PyTorch to #WebAI using…

X AI KOLs Timeline · 2026-07-09 Cached

Google announces LiteRT.js, a high-performance JavaScript runtime for running AI models directly in the browser using WebAssembly and hardware acceleration, as an evolution from TensorFlow.js.

0 favorites 0 likes
#on-device-inference

@googlegemma: “Agentic kernel optimization is the future of on-device inference” @xenovacom used Fable 5 to write kernels that pushed…

X AI KOLs Timeline · 2026-07-01 Cached

Xenova used Fable 5 to write optimized kernels achieving 255 tokens per second for Gemma 4 on WebGPU with M4, demonstrating agentic kernel optimization for on-device inference.

0 favorites 0 likes
#on-device-inference

New MLX LM Server From Apple

Reddit r/LocalLLaMA · 2026-06-09 Cached

Apple's MLX team introduces MLX LM Server, a tool for running AI agent workflows fully locally on Mac, supporting continuous batching, distributed inference, and M5 neural acceleration, with no need for cloud or API keys.

0 favorites 0 likes
#on-device-inference

Ran gemma 4 12b on my 3090 yesterday and I think the local model game just changed

Reddit r/artificial · 2026-06-04

A user reports running Google's Gemma 4 12B model locally on a single RTX 3090 via GGUF quantization, finding strong performance including real 256k context, multimodal capabilities, and function calling that outperforms larger 70B models for coding tasks.

0 favorites 0 likes
#on-device-inference

When Cloud Agents Meet Device Agents: Lessons from Hybrid Multi-Agent Systems

Hugging Face Daily Papers · 2026-05-28 Cached

This paper systematically studies hybrid multi-agent systems combining cloud-based LLMs and on-device SLMs, revealing task-dependent optimal architectures and challenging the assumption that more frontier compute always improves performance.

0 favorites 0 likes
#on-device-inference

MobileExplorer: Accelerating On-Device Inference for Mobile GUI Agents via Online Exploration

arXiv cs.AI · 2026-05-27 Cached

MobileExplorer is a new framework that accelerates on-device inference for mobile GUI agents by performing lightweight parallel exploration of UI elements during model inference, reducing reasoning steps and latency by 23% while maintaining or improving task success rates.

0 favorites 0 likes
#on-device-inference

ExecuTorch -- A Unified PyTorch Solution to Run AI Models On-Device

arXiv cs.LG · 2026-05-12 Cached

This article introduces ExecuTorch, a unified PyTorch-native deployment framework designed to run AI models on diverse edge devices without requiring model conversion or reimplementation.

0 favorites 0 likes
#on-device-inference

Local AI needs to be the norm

Hacker News Top · 2026-05-10 Cached

The article argues against relying on cloud-hosted AI APIs due to privacy and reliability concerns, advocating for on-device AI processing as demonstrated by a native iOS app using Apple's local model APIs.

0 favorites 0 likes
#on-device-inference

Do you think edge AI ends up mattering more for autonomy, robotics, or local private inference?

Reddit r/artificial · 2026-05-08

A discussion post exploring where edge AI will have the greatest impact: autonomy and robotics, low-power vision systems, private local LLMs, or bandwidth-constrained industrial deployments.

0 favorites 0 likes
← Back to home

Submit Feedback