offline-inference

Tag

Cards List
#offline-inference

React Native ExecuTorch now runs Gemma 4 (Vulkan and MLX accelerated)

Reddit r/LocalLLaMA · 2026-06-15

The react-native-executorch library now integrates Google's Gemma 4 model, enabling fully offline, GPU-accelerated inference in React Native apps using Vulkan on Android and MLX on Apple Silicon.

0 favorites 0 likes
#offline-inference

@danveloper: https://x.com/danveloper/status/2064387956387758206

X AI KOLs Timeline · 2026-06-09 Cached

A developer ran DeepSeek-V4-Flash on a Raspberry Pi 5 by streaming model weights from an NVMe SSD, achieving 1.3 tokens/second at 8 watts, demonstrating the feasibility of frontier-adjacent open-weight models on low-cost, offline hardware.

0 favorites 0 likes
#offline-inference

Gemma 4 running fully offline on WebGPU with Transformers.js, controlling Reachy Mini over WebSerial.

Reddit r/LocalLLaMA · 2026-05-11

Demonstrates running Gemma 4 offline in the browser using WebGPU and Transformers.js to control a Reachy Mini robot via WebSerial.

0 favorites 0 likes
#offline-inference

What impedes apps using AI to make the user’s device the server running a local LLM?

Reddit r/singularity · 2026-04-22

A user reflects on why more apps don’t run local LLMs directly on phones, noting Gemma 2-4B models already work offline and could eliminate server costs while maintaining near-GPT-4o quality.

0 favorites 0 likes
← Back to home

Submit Feedback