offline-inference

Tag

Cards List
#offline-inference

PTC-Decoder: Towards Intelligent SLMs on Offline Resource-Constrained Edge Devices

arXiv cs.AI ↗ · 2d ago Cached

PTC-Decoder is a training-free, plug-and-play framework that improves small language models on offline, resource-constrained edge devices by enforcing plan adherence through token-level constraints, enhancing step-level reliability in multi-step agent tasks.

0 favorites 0 likes
#offline-inference

@seclink: @grok Summarize it: what novel, open-source capabilities have emerged this time.

X AI KOLs Timeline ↗ · 6d ago Cached

Google Antigravity SDK introduces the capability to run open-source AI models like Gemma 4 completely offline on local GPUs, ensuring zero API costs and total data privacy.

0 favorites 0 likes
#offline-inference

SHADOW 50M: a 19.8 MB model that computes exactly and remembers from disk [P]

Reddit r/MachineLearning ↗ · 2026-09-14

SHADOW 50M is a compact 44M-parameter AI model with ternary weights and fixed 512-bit codes, trained from scratch on 45B tokens, enabling offline inference at high speed on consumer hardware.

1 favorites 1 likes
#offline-inference

@AlmustyFX: This is the kind of local AI test that actually matters. A 7.9B MoE running at 152 tok/s on an M4 Pro with 64GB unified…

X AI KOLs Timeline ↗ · 2026-08-29 Cached

A 7.9B MoE model runs at 152 tokens per second on an M4 Pro with 64GB unified memory, enabling offline processing of sensitive contract data and demonstrating the practical use of local AI.

0 favorites 0 likes
#offline-inference

React Native ExecuTorch now runs Gemma 4 (Vulkan and MLX accelerated)

Reddit r/LocalLLaMA ↗ · 2026-06-15

The react-native-executorch library now integrates Google's Gemma 4 model, enabling fully offline, GPU-accelerated inference in React Native apps using Vulkan on Android and MLX on Apple Silicon.

0 favorites 0 likes
#offline-inference

@danveloper: https://x.com/danveloper/status/2064387956387758206

X AI KOLs Timeline ↗ · 2026-06-09 Cached

A developer ran DeepSeek-V4-Flash on a Raspberry Pi 5 by streaming model weights from an NVMe SSD, achieving 1.3 tokens/second at 8 watts, demonstrating the feasibility of frontier-adjacent open-weight models on low-cost, offline hardware.

0 favorites 0 likes
#offline-inference

Gemma 4 running fully offline on WebGPU with Transformers.js, controlling Reachy Mini over WebSerial.

Reddit r/LocalLLaMA ↗ · 2026-05-11

Demonstrates running Gemma 4 offline in the browser using WebGPU and Transformers.js to control a Reachy Mini robot via WebSerial.

0 favorites 0 likes
#offline-inference

What impedes apps using AI to make the user’s device the server running a local LLM?

Reddit r/singularity ↗ · 2026-04-22

A user reflects on why more apps don’t run local LLMs directly on phones, noting Gemma 2-4B models already work offline and could eliminate server costs while maintaining near-GPT-4o quality.

0 favorites 0 likes
← Back to home

Submit Feedback