Open Dungeon: local roleplay with Gemma 4 QAT + inline Uncen-FLUX images, running at full 256K context under 8GB RAM (OS)
Summary
An open-source local AI dungeon app using Gemma 4 and FLUX for text and image generation, fully private and runs under 8GB RAM.
Similar Articles
@UnslothAI: Gemma 4 12B can now run locally on just 8GB RAM via Dynamic GGUFs. Google's new model, Gemma 4 12B Unified supports ima…
Gemma 4 12B, Google's multimodal open model supporting image, audio, and 256K context, can now run locally on just 8GB RAM via Unsloth's Dynamic GGUFs, enabling local training and inference through Unsloth Studio.
@analogalok: Run Gemma 4 26B MoE on 8GB VRAM with 250k context at 20+ tokens/sec If you own any 8GB VRAM graphics card, stop what yo…
Alok demonstrates running Gemma 4 26B MoE on 8GB VRAM using Unsloth's QAT quant and the -cmoe flag in llama.cpp, achieving 20 tokens/sec with 250k context, marking a major milestone for budget local AI.
yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF
A focused fine-tune of Gemma 4 12B for coding, distilled from chain-of-thought data (Composer 2.5 and Fable 5) and quantized to GGUF for local, offline use with minimal VRAM requirements.
I made Warrior Quest, a local LLM-powered dark-fantasy RPG where the model only plays NPCs and the actual game state stays deterministic
Warrior Quest is a local LLM-powered dark-fantasy RPG demo on Steam where the AI model is used only for NPC emulation while the game state remains deterministic, requiring a GPU with 8GB VRAM.
Experiment: autonomous NPCs powered by Gemma 4 E2B in the browser
An experiment demonstrating autonomous NPCs in the browser powered by Gemma 4 and E2B.