4GB "Gemini Nano" model GGUF anyone?
Summary
A user inquires about the specific identity of a ~4GB AI model (likely Gemini Nano) silently downloaded by Chrome for on-device features, and requests a GGUF version for local execution via llama.cpp.
Similar Articles
@_ar9av: So apparently Google ships a Gemini Nano 4B LLM (context limit : 9216 tokens) baked into chrome I tried to expose it ou…
A developer discovered Google's Gemini Nano 4B LLM built into Chrome and created an OpenAI-compatible API wrapper for local use, eliminating the need for API keys or external calls.
Jiunsong/supergemma4-26b-uncensored-gguf-v2
SuperGemma4-26B-Uncensored-Fast GGUF v2 is a quantized, locally-runnable variant of Google's Gemma-4-26B model optimized for Apple Silicon, offering faster inference speeds and less-censored chat behavior while maintaining practical performance on general tasks.
guess what? if you are a chrome user, technically you are localllama member!
Google Chrome is silently installing a 4 GB Gemini Nano AI model on user devices without explicit consent or opt-out UI, raising significant privacy, legal, and environmental concerns.
Run Chrome’s tiny Gemma4 (aka Gemini Nano) directly on PC without GPU
A developer created a Chrome extension called Dobby that runs Google's Gemma4 (Gemini Nano) locally on PC without needing a GPU, requiring only Chrome and 16GB RAM. The extension provides a simple interface to interact with the model for tasks like spell checking or summarizing.
@UnslothAI: Gemma 4 12B can now run locally on just 8GB RAM via Dynamic GGUFs. Google's new model, Gemma 4 12B Unified supports ima…
Gemma 4 12B, Google's multimodal open model supporting image, audio, and 256K context, can now run locally on just 8GB RAM via Unsloth's Dynamic GGUFs, enabling local training and inference through Unsloth Studio.