NVIDIA and Google collaborate to optimize Gemma 4 models for local deployment across RTX GPUs, DGX Spark, and Jetson devices, enabling efficient on-device agentic AI with support for reasoning, coding, multimodal capabilities, and 35+ languages.
<div id="bsf_rt_marker"></div><p><span data-contrast="none">Open models are driving a new wave of on-device AI, extending innovation beyond the cloud to everyday devices. As these models advance, their value increasingly depends on access to local, real-time context that can turn meaningful insights into action.</span><span data-ccp-props="{"335559739":0}"> </span></p>
<p><span data-contrast="none">Designed for this shift, Google’s latest additions to the </span>Gemma 4 family <span data-contrast="none">introduce a class of small, fast and omni-capable models built for efficient local execution across a wide range of devices. </span><span data-ccp-props="{"335559739":0}"> </span></p>
<p>Google <span data-contrast="none">and NVIDIA have collaborated to optimize Gemma 4</span><b><span data-contrast="none"> </span></b><span data-contrast="none">for NVIDIA GPUs, enabling efficient performance across a range of systems — from data center deployments to NVIDIA RTX-powered PCs and workstations, the <a target="_blank" href="https://www.nvidia.com/en-us/products/workstations/dgx-spark/">NVIDIA DGX Spark</a> personal AI supercomputer and <a target="_blank" href="https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-nano/product-development/">NVIDIA Jetson Orin Nano</a> edge AI modules.</span></p>
<h2><b><span data-contrast="none">Gemma 4: Compact Models Optimized for NVIDIA GPUs</span></b><span data-ccp-props="{"335559739":0}"> </span></h2>
<p><span data-contrast="none">The latest additions to the </span>Gemma 4 family of open models<span data-contrast="none">—</span><span data-contrast="none"> spanning E2B, E4B, 26B and 31B variants </span><span data-contrast="none">—</span><span data-contrast="none"> are designed for efficient deployment from edge devices to high-performance GPUs. </span><span data-ccp-props="{"335559739":0}"> </span></p>
<figure id="attachment_92036" aria-describedby="caption-attachment-92036" style="width: 1149px" class="wp-caption aligncenter"><a href="https://blogs.nvidia.com/wp-content/uploads/2026/04/gemma-4-perf-chart-desktop-light-1.png"><img loading="lazy" decoding="async" class="size-full wp-image-92036" src="https://blogs.nvidia.com/wp-content/uploads/2026/04/gemma-4-perf-chart-desktop-light-1.png" alt="" width="1149" height="489" srcset="https://blogs.nvidia.com/wp-content/uploads/2026/04/gemma-4-perf-chart-desktop-light-1.png 1149w, https://blogs.nvidia.com/wp-content/uploads/2026/04/gemma-4-perf-chart-desktop-light-1-960x409.png 960w, https://blogs.nvidia.com/wp-content/uploads/2026/04/gemma-4-perf-chart-desktop-light-1-630x268.png 630w" sizes="auto, (max-width: 1149px) 100vw, 1149px" /></a><figcaption id="caption-attachment-92036" class="wp-caption-text">All configurations measured using Q4_K_M quantizations BS = 1, ISL = 4096 and OSL = 128 on NVIDIA GeForce RTX 5090 and Mac M3 Ultra desktops. Token generation throughput measured on llama.cpp b7789, using the llama-bench tool.</figcaption></figure>
<p><span data-contrast="none">This new generation of compact models supports a range of tasks, including:</span><span data-ccp-props="{"335559739":0}"> </span></p>
<ul>
<li><b><span data-contrast="none">Reasoning: </span></b><span data-contrast="none">Strong performance on complex problem-solving tasks. </span><span data-ccp-props="{"335559739":0}"> </span></li>
<li><b><span data-contrast="auto">Coding: </span></b><span data-contrast="auto">Code generation and debugging for developer workflows. </span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":240,"335559739":240}"> </span></li>
<li><b><span data-contrast="auto">Agents: </span></b><span data-contrast="auto">Native support for structured tool use (function calling). </span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":240,"335559739":240}"> </span></li>
<li><b><span data-contrast="auto">Vision, Video and Audio Capabilities: </span></b><span data-contrast="auto">E</span><span data-contrast="auto">nables rich multimodal interactions for object recognition, automated speech recognition, and document or video intelligence.</span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":240,"335559739":240}"> </span></li>
<li><b><span data-contrast="auto">Interleaved Multimodal Input: </span></b><span data-contrast="auto">M</span><span data-contrast="auto">ix text and images in any order within a single prompt. </span><span data-ccp-props="{"134233117":false,"134233118":false,"335559738":240,"335559739":240}"> </span></li>
<li><b><span data-contrast="auto">Multilingual: </span></b><span data-contrast="auto">Out-of-the-box support for 35+ languages, pretrained on 140+ languages.</span><span data-ccp-props="{}"> </span></li>
</ul>
<p><span data-contrast="none">The </span>E2B and E4B models<span data-contrast="none"> are built for ultraefficient, low-latency inference at the edge, running completely offline with near-zero latency across many devices including Jetson Nano modules. </span></p>
<p><span data-contrast="none">The </span>26B and 31B models<span data-contrast="none">are designed for high-performance reasoning and developer-centric workflows, making them well suited for agentic AI. Optimized to deliver state-of-the-art, accessible reasoning, these models run efficiently on NVIDIA RTX GPUs and DGX Spark — powering development environments, coding assistants and agent-driven workflows. </span><span data-ccp-props="{"335559739":0}"> </span></p>
<p><span data-contrast="none">As local agentic AI continues to gain momentum, applications like </span>OpenClaw<span data-contrast="none"> are enabling always-on AI assistants on RTX PCs, workstations and DGX Spark. The latest Gemma 4 models are compatible with OpenClaw, allowing users to build capable local agents that draw context from personal files, applications and workflows to automate tasks. Learn how to run </span><a target="_blank" href="https://www.nvidia.com/en-us/geforce/news/open-claw-rtx-gpu-dgx-spark-guide/"><span data-contrast="none">OpenClaw for free on RTX GPUs and DGX Spark</span></a><span data-contrast="none"> or using the </span><a target="_blank" href="https://build.nvidia.com/spark/openclaw"><span data-contrast="none">DGX Spark OpenClaw playbook</span></a><span data-contrast="auto">.</span><span data-ccp-props="{"335559739":0}"> </span></p>
<p><span class="NormalTextRun CommentStart CommentHighlightPipeRest CommentHighlightRest SCXW107558427 BCX0">Check out the <a target="_blank" href="https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/">Google DeepMind announcement blog</a> to l</span><span class="NormalTextRun CommentHighlightRest SCXW107558427 BCX0">earn more about the</span><span class="NormalTextRun CommentHighlightRest SCXW107558427 BCX0"> </span><span class="NormalTextRun CommentHighlightRest SCXW107558427 BCX0">latest additions to </span><span class="NormalTextRun CommentHighlightRest SCXW107558427 BCX0">Gemma 4 </span><span class="NormalTextRun CommentHighlightRest SCXW107558427 BCX0">family.</span></p>
<h2><b><span data-contrast="none">Getting Started: Gemma 4 on RTX GPUs and DGX Spark</span></b><span data-ccp-props="{"335559739":0}"> </span></h2>
<p><span data-contrast="none">NVIDIA has collaborated with Ollama and llama.cpp to provide the best local deployment experience for each of the Gemma 4 models. </span><span data-ccp-props="{"335559739":0}"> </span></p>
<p><span data-contrast="none">To use Gemma 4 locally, users can </span><span data-contrast="none">download Ollama</span><span data-contrast="none"> to run Gemma 4 models </span><span data-contrast="none">or</span><span data-contrast="none"> install </span><span data-contrast="none">llama.cpp</span><span data-contrast="none"> and pair it with the Gemma 4 GGUF Hugging Face checkpoint. </span><span data-contrast="auto">Additionally, </span><span data-contrast="none">Unsloth provides day-one support with optimized and quantized models for efficient local fine-tuning and deployment via Unsloth Studio. Start </span><span data-contrast="auto">running and </span><span data-contrast="none">fine-tuning</span><span data-contrast="auto"> Gemma 4 in Unsloth Studio today.</span><span data-ccp-props="{"335559739":0}"> </span></p>
<p><span data-contrast="none">Running open models like the Gemma 4 family on NVIDIA GPUs achieves optimal performance because NVIDIA Tensor Cores accelerate AI inference workloads to deliver higher throughput and lower latency for local execution. Plus, the CUDA software stack ensures broad compatibility across leading frameworks and tools, enabling new models to run efficiently from day one. </span><span data-ccp-props="{"335559739":0}"> </span></p>
<p><span data-contrast="none">This combination allows open models like Gemma 4 to scale across a wide range of systems — from Jetson Orin Nano at the edge to RTX PCs, workstations and DGX Spark — without requiring extensive optimization.</span><span data-ccp-props="{"335559739":0}"> </span></p>
<p><span data-contrast="none">Check out </span><span data-contrast="none">the </span><a target="_blank" href="https://developer.nvidia.com/blog/bringing-ai-closer-to-the-edge-and-on-device-with-gemma-4/"><span data-contrast="none">NVIDIA technical blog</span></a><span data-contrast="none"> </span><span data-contrast="none">for more details on how to get started with Gemma 4 on NVIDIA GPUs and learn more about</span><span data-contrast="none"> NVIDIA’s work on </span><a href="https://blogs.nvidia.com/blog/ai-future-open-and-proprietary/"><span data-contrast="none">open models</span></a><span data-contrast="none">.</span><span data-ccp-props="{"335559739":0}"> </span></p>
<h2><b><span data-contrast="none">#ICYMI: The Latest Updates for RTX AI PCs</span></b><span data-ccp-props="{"335559739":0}"> </span></h2>
<p><span data-contrast="none"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2728.png" alt="✨" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Catch up on </span><a href="https://blogs.nvidia.com/blog/rtx-ai-garage-gtc-2026-nemoclaw"><span data-contrast="none">RTX AI Garage</span></a><span data-contrast="none"> blogs for a host of agentic AI announcements from NVIDIA GTC, such as new open models for local agents. These models include NVIDIA Nemotron 3 Nano 4B and Nemotron 3 Super 120B, and optimizations for Qwen 3.5 and Mistral Small 4.</span><span data-ccp-props="{"335559739":0}"> </span></p>
<p><span data-contrast="none"> NVIDIA recently introduced </span><a target="_blank" href="https://nvidianews.nvidia.com/news/nvidia-announces-nemoclaw"><span data-contrast="none">NVIDIA NemoClaw,</span></a><span data-contrast="none"> an open source stack that optimizes OpenClaw experiences on NVIDIA devices by increasing security and supporting local models. </span><span data-ccp-props="{"335559739":0}"> </span></p>
<p><b><span data-contrast="none"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f680.png" alt="🚀" class="wp-smiley" style="height: 1em; max-height: 1em;" /></span></b><b><span data-contrast="none"> </span></b><a target="_blank" href="https://accomplish.ai/"><span data-contrast="none">Accomplish.ai</span></a><span data-contrast="none"> announced Accomplish FREE, a no-cost version of its open source desktop AI agent with built-in models. It harnesses NVIDIA GPUs to run open weight models locally, while a hybrid router dynamically balances workloads between local RTX hardware and the cloud — enabling fast, private, zero-configuration execution without requiring an application programming interface key.</span><span data-ccp-props="{"335559739":0}"> </span></p>
<p><i><span data-contrast="none">Plug in to NVIDIA AI PC on </span></i><a target="_blank" href="https://www.facebook.com/NVIDIA.AI.PC/"><i><span data-contrast="none">Facebook</span></i></a><i><span data-contrast="none">, </span></i><a target="_blank" href="https://www.instagram.com/nvidia.ai.pc/"><i><span data-contrast="none">Instagram</span></i></a><i><span data-contrast="none">, </span></i><a target="_blank" href="https://www.tiktok.com/@nvidia_ai_pc"><i><span data-contrast="none">TikTok</span></i></a><i><span data-contrast="none"> and </span></i><a target="_blank" href="https://x.com/NVIDIA_AI_PC"><i><span data-contrast="none">X</span></i></a><i><span data-contrast="none"> — and stay informed by subscribing to the </span></i><a target="_blank" href="https://www.nvidia.com/en-us/ai-on-rtx/?modal=subscribe-ai"><i><span data-contrast="none">RTX AI PC newsletter</span></i></a><i><span data-contrast="none">.</span></i><span data-ccp-props="{"335559739":0}"> </span></p>
<p><i><span data-contrast="none">Follow NVIDIA Workstation on </span></i><a target="_blank" href="https://www.linkedin.com/showcase/3761136/"><i><span data-contrast="none">LinkedIn</span></i></a><i><span data-contrast="none"> and </span></i><a target="_blank" href="https://x.com/NVIDIAworkstatn"><i><span data-contrast="none">X</span></i></a><i><span data-contrast="none">. </span></i><span data-ccp-props="{"335559739":0}"> </span></p>
# From RTX to Spark: NVIDIA Accelerates Gemma 4 for Local Agentic AI
Source: [https://blogs.nvidia.com/blog/rtx-ai-garage-open-models-google-gemma-4/](https://blogs.nvidia.com/blog/rtx-ai-garage-open-models-google-gemma-4/)
Open models are driving a new wave of on\-device AI, extending innovation beyond the cloud to everyday devices\. As these models advance, their value increasingly depends on access to local, real\-time context that can turn meaningful insights into action\.
Designed for this shift, Google’s latest additions to theGemma 4 familyintroduce a class of small, fast and omni\-capable models built for efficient local execution across a wide range of devices\.
Googleand NVIDIA have collaborated to optimize Gemma 4for NVIDIA GPUs, enabling efficient performance across a range of systems — from data center deployments to NVIDIA RTX\-powered PCs and workstations, the[NVIDIA DGX Spark](https://www.nvidia.com/en-us/products/workstations/dgx-spark/)personal AI supercomputer and[NVIDIA Jetson Orin Nano](https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-nano/product-development/)edge AI modules\.
## **Gemma 4: Compact Models Optimized for NVIDIA GPUs**
The latest additions to theGemma 4 family of open models—spanning E2B, E4B, 26B and 31B variants—are designed for efficient deployment from edge devices to high\-performance GPUs\.
[](https://blogs.nvidia.com/wp-content/uploads/2026/04/gemma-4-perf-chart-desktop-light-1.png)All configurations measured using Q4\_K\_M quantizations BS = 1, ISL = 4096 and OSL = 128 on NVIDIA GeForce RTX 5090 and Mac M3 Ultra desktops\. Token generation throughput measured on llama\.cpp b7789, using the llama\-bench tool\.This new generation of compact models supports a range of tasks, including:
- **Reasoning:**Strong performance on complex problem\-solving tasks\.
- **Coding:**Code generation and debugging for developer workflows\.
- **Agents:**Native support for structured tool use \(function calling\)\.
- **Vision, Video and Audio Capabilities:**Enables rich multimodal interactions for object recognition, automated speech recognition, and document or video intelligence\.
- **Interleaved Multimodal Input:**Mix text and images in any order within a single prompt\.
- **Multilingual:**Out\-of\-the\-box support for 35\+ languages, pretrained on 140\+ languages\.
TheE2B and E4B modelsare built for ultraefficient, low\-latency inference at the edge, running completely offline with near\-zero latency across many devices including Jetson Nano modules\.
The26B and 31B modelsare designed for high\-performance reasoning and developer\-centric workflows, making them well suited for agentic AI\. Optimized to deliver state\-of\-the\-art, accessible reasoning, these models run efficiently on NVIDIA RTX GPUs and DGX Spark — powering development environments, coding assistants and agent\-driven workflows\.
As local agentic AI continues to gain momentum, applications likeOpenClaware enabling always\-on AI assistants on RTX PCs, workstations and DGX Spark\. The latest Gemma 4 models are compatible with OpenClaw, allowing users to build capable local agents that draw context from personal files, applications and workflows to automate tasks\. Learn how to run[OpenClaw for free on RTX GPUs and DGX Spark](https://www.nvidia.com/en-us/geforce/news/open-claw-rtx-gpu-dgx-spark-guide/)or using the[DGX Spark OpenClaw playbook](https://build.nvidia.com/spark/openclaw)\.
## **Getting Started: Gemma 4 on RTX GPUs and DGX Spark**
NVIDIA has collaborated with Ollama and llama\.cpp to provide the best local deployment experience for each of the Gemma 4 models\.
To use Gemma 4 locally, users candownload Ollamato run Gemma 4 modelsorinstallllama\.cppand pair it with the Gemma 4 GGUF Hugging Face checkpoint\.Additionally,Unsloth provides day\-one support with optimized and quantized models for efficient local fine\-tuning and deployment via Unsloth Studio\. Startrunning andfine\-tuningGemma 4 in Unsloth Studio today\.
Running open models like the Gemma 4 family on NVIDIA GPUs achieves optimal performance because NVIDIA Tensor Cores accelerate AI inference workloads to deliver higher throughput and lower latency for local execution\. Plus, the CUDA software stack ensures broad compatibility across leading frameworks and tools, enabling new models to run efficiently from day one\.
This combination allows open models like Gemma 4 to scale across a wide range of systems — from Jetson Orin Nano at the edge to RTX PCs, workstations and DGX Spark — without requiring extensive optimization\.
Check outthe[NVIDIA technical blog](https://developer.nvidia.com/blog/bringing-ai-closer-to-the-edge-and-on-device-with-gemma-4/)for more details on how to get started with Gemma 4 on NVIDIA GPUs and learn more aboutNVIDIA’s work on[open models](https://blogs.nvidia.com/blog/ai-future-open-and-proprietary/)\.
## **\#ICYMI: The Latest Updates for RTX AI PCs**
✨ Catch up on[RTX AI Garage](https://blogs.nvidia.com/blog/rtx-ai-garage-gtc-2026-nemoclaw)blogs for a host of agentic AI announcements from NVIDIA GTC, such as new open models for local agents\. These models include NVIDIA Nemotron 3 Nano 4B and Nemotron 3 Super 120B, and optimizations for Qwen 3\.5 and Mistral Small 4\.
NVIDIA recently introduced[NVIDIA NemoClaw,](https://nvidianews.nvidia.com/news/nvidia-announces-nemoclaw)an open source stack that optimizes OpenClaw experiences on NVIDIA devices by increasing security and supporting local models\.
**🚀**[Accomplish\.ai](https://accomplish.ai/)announced Accomplish FREE, a no\-cost version of its open source desktop AI agent with built\-in models\. It harnesses NVIDIA GPUs to run open weight models locally, while a hybrid router dynamically balances workloads between local RTX hardware and the cloud — enabling fast, private, zero\-configuration execution without requiring an application programming interface key\.
*Plug in to NVIDIA AI PC on*[*Facebook*](https://www.facebook.com/NVIDIA.AI.PC/)*,*[*Instagram*](https://www.instagram.com/nvidia.ai.pc/)*,*[*TikTok*](https://www.tiktok.com/@nvidia_ai_pc)*and*[*X*](https://x.com/NVIDIA_AI_PC)*— and stay informed by subscribing to the*[*RTX AI PC newsletter*](https://www.nvidia.com/en-us/ai-on-rtx/?modal=subscribe-ai)*\.*
*Follow NVIDIA Workstation on*[*LinkedIn*](https://www.linkedin.com/showcase/3761136/)*and*[*X*](https://x.com/NVIDIAworkstatn)*\.*
NVIDIA optimizes Google DeepMind's DiffusionGemma, an open model that generates text in parallel 256-token blocks, achieving up to 4x faster performance on local RTX GPUs, DGX Spark, and DGX Station systems.
NVIDIA announced RTX Spark PCs and a wave of updates to enable local AI agents across RTX and DGX ecosystems, including the OpenShell runtime coming to Windows, NemoClaw expansion, performance improvements, and integrations with Adobe and H Company.
Google announces the availability of Gemma 4 12B on laptops via Google AI Edge, enabling local, agentic, and multimodal workflows with tools like AI Edge Gallery and Eloquent.
Google released Gemma 4, an open-source AI model optimized for local execution on standard laptops, offering 3x faster performance and a 256k context window for free under an Apache 2.0 license.
Google announced Gemma 4 E2B optimized for the Pixel 10's TPU, enabling on-device multimodal AI capabilities like offline chat, image recognition, and audio transcription.