Gemma 4 VLA Demo on Jetson Orin Nano Super
Summary
NVIDIA and Hugging Face publish a hands-on demo showing Gemma 4 running as a vision-language-action model entirely on the Jetson Orin Nano Super, using local STT/TTS and webcam input.
View Cached Full Text
Cached at: 04/22/26, 08:23 PM
Gemma 4 VLA Demo on Jetson Orin Nano Super
Source: https://huggingface.co/blog/nvidia/gemma4 Back to Articles
https://huggingface.co/login?next=%2Fblog%2Fnvidia%2Fgemma4- ![]()
- Get the code
- Hardware
- Step 1: System packages
- Step 2: Python environment
- Step 3: Free up RAM (optional but recommended)- Add some swap - Kill memory hogs - Still tight on RAM?
- Step 4: Serve Gemma 4- Build llama.cpp - Download the model and vision projector - Start the server - Verify it’s up
- Step 5: Find your mic, speaker, and webcam- Microphone - Speaker - Webcam - Quick test
- Step 6: Run the demo- Changing the voice
- How it works
- Troubleshooting
- Environment variables
- Bonus: just want to try Gemma 4 in text mode?
Talk to Gemma 4, and she’ll decide on her own if she needs to look through the webcam to answer you. All running locally on a Jetson Orin Nano Super.
You speak → Parakeet STT → Gemma 4 → [Webcam if needed] → Kokoro TTS → Speaker
Press SPACE to record, SPACE again to stop. This is a simple VLA: the model decides on its own whether to act based on the context of what you asked, no keyword triggers, no hardcoded logic. If your question needs Gemma to open her eyes, she’ll decide to take a photo, interpret it, and answer you with that context in mind. She’s not describing the picture, she’s answering your actual question using what she saw.
And honestly? It’s pretty impressive that this runs on a Jetson Orin Nano. :)
https://huggingface.co/blog/nvidia/gemma4#get-the-codeGet the code
The full script for this tutorial lives on GitHub, in myGoogle\_Gemmarepo next to the Gemma 2 demos:
👉github.com/asierarranz/Google_Gemma
Grab it with either of these (pick one):
# Option 1: clone the whole repo
git clone https://github.com/asierarranz/Google_Gemma.git
cd Google_Gemma/Gemma4
# Option 2: just download the script
wget https://raw.githubusercontent.com/asierarranz/Google_Gemma/main/Gemma4/Gemma4_vla.py
That single file (Gemma4\_vla\.py) is all you need. It pulls the STT/TTS models and voice assets from Hugging Face on first run.
https://huggingface.co/blog/nvidia/gemma4#hardwareHardware
What we used:
- NVIDIA Jetson Orin Nano Super(8 GB)
- Logitech C920 webcam (mic built in)
- USB speaker
- USB keyboard (to press SPACE)
Not tied to these exact devices, any webcam, USB mic, and USB speaker that Linux sees should work.
https://huggingface.co/blog/nvidia/gemma4#step-1-system-packagesStep 1: System packages
Fresh Jetson, let’s install the basics:
sudo apt update
sudo apt install -y \
git build-essential cmake curl wget pkg-config \
python3-pip python3-venv python3-dev \
alsa-utils pulseaudio-utils v4l-utils psmisc \
ffmpeg libsndfile1
build\-essentialandcmakeare only needed if you go the native llama.cpp route (Option A in Step 4). The rest is for audio, webcam, and Python.
https://huggingface.co/blog/nvidia/gemma4#step-2-python-environmentStep 2: Python environment
python3 -m venv .venv
source .venv/bin/activate
pip install --upgrade pip
pip install opencv-python-headless onnx_asr kokoro-onnx soundfile huggingface-hub numpy
https://huggingface.co/blog/nvidia/gemma4#step-3-free-up-ram-optional-but-recommendedStep 3: Free up RAM (optional but recommended)
Heads up: this step may not be strictly necessary. But we’re pushing this 8 GB board pretty hard with a fairly capable model, so giving ourselves some headroom makes the whole experience smoother, especially if you’ve been playing with Docker or other heavy stuff before this.
These are just the commands that worked nicely for me. Use them if they help.
https://huggingface.co/blog/nvidia/gemma4#add-some-swapAdd some swap
Swap won’t speed up inference, but it acts as a safety net during model loading so you don’t get OOM-killed at the worst moment.
sudo fallocate -l 8G /swapfile
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile
echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab
https://huggingface.co/blog/nvidia/gemma4#kill-memory-hogsKill memory hogs
sudo systemctl stop docker 2>/dev/null || true
sudo systemctl stop containerd 2>/dev/null || true
pkill -f tracker-miner-fs-3 || true
pkill -f gnome-software || true
free -h
Close browser tabs, IDE windows, anything you don’t need. Every MB counts.
If you’re going with the Docker route in Step 4, obviously don’t stop Docker here, you’ll need it. Still kill the rest though.
https://huggingface.co/blog/nvidia/gemma4#still-tight-on-ramStill tight on RAM?
From our tests,Q4\_K\_M(native build) andQ4\_K\_S(Docker) run comfortably on the 8 GB board once you’ve done the cleanup above. But if you’ve got other stuff you can’t kill and memory is still tight, you can drop one step down to aQ3quant, same model, a bit less smart, noticeably lighter. Just swap the filename in Step 4:
gemma-4-E2B-it-Q3_K_M.gguf # instead of Q4_K_M
Honestly though, stick withQ4\_K\_Mif you can. It’s the sweet spot.
https://huggingface.co/blog/nvidia/gemma4#step-4-serve-gemma-4Step 4: Serve Gemma 4
You need a runningllama\-serverwith Gemma 4 before launching the demo. We’ll buildllama\.cppnatively on the Jetson, it gives the best performance and full control over the vision projector that the VLA demo needs.
https://huggingface.co/blog/nvidia/gemma4#build-llamacppBuild llama.cpp
cd ~
git clone https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
cmake -B build \
-DGGML_CUDA=ON \
-DCMAKE_CUDA_ARCHITECTURES="87" \
-DGGML_NATIVE=ON \
-DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release -j4
https://huggingface.co/blog/nvidia/gemma4#download-the-model-and-vision-projectorDownload the model and vision projector
mkdir -p ~/models && cd ~/models
wget -O gemma-4-E2B-it-Q4_K_M.gguf \
https://huggingface.co/unsloth/gemma-4-E2B-it-GGUF/resolve/main/gemma-4-E2B-it-Q4_K_M.gguf
wget -O mmproj-gemma4-e2b-f16.gguf \
https://huggingface.co/ggml-org/gemma-4-E2B-it-GGUF/resolve/main/mmproj-gemma4-e2b-f16.gguf
Themmprojfile is the vision projector. Without it Gemma can’t see, so don’t skip it.
https://huggingface.co/blog/nvidia/gemma4#start-the-serverStart the server
~/llama.cpp/build/bin/llama-server \
-m ~/models/gemma-4-E2B-it-Q4_K_M.gguf \
--mmproj ~/models/mmproj-gemma4-e2b-f16.gguf \
-c 2048 \
--image-min-tokens 70 --image-max-tokens 70 \
--ubatch-size 512 --batch-size 512 \
--host 0.0.0.0 --port 8080 \
-ngl 99 --flash-attn on \
--no-mmproj-offload --jinja -np 1
One flag worth mentioning:\-ngl 99tellsllama\-serverto push all the model’s layers onto the GPU (99 is just “as many as the model has”). If you ever run into memory issues, you can lower that number to offload fewer layers to GPU and the rest to CPU. For this setup though, all layers on GPU should work fine.
https://huggingface.co/blog/nvidia/gemma4#verify-its-upVerify it’s up
From another terminal:
curl -s http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"gemma4","messages":[{"role":"user","content":"Hi!"}],"max_tokens":32}' \
| python3 -m json.tool
If you get JSON back, you’re good.
https://huggingface.co/blog/nvidia/gemma4#step-5-find-your-mic-speaker-and-webcamStep 5: Find your mic, speaker, and webcam
https://huggingface.co/blog/nvidia/gemma4#microphoneMicrophone
arecord -l
Look for your USB mic. In our case the C920 showed up asplughw:3,0.
https://huggingface.co/blog/nvidia/gemma4#speakerSpeaker
pactl list short sinks
This lists your PulseAudio sinks. Pick the one that matches your speaker, it’ll be a long ugly name likealsa\_output\.usb\-\.\.\.. In my case it wasalsa\_output\.usb\-Generic\_USB2\.0\_Device\_20130100ph0\-00\.analog\-stereo, but yours will be different.
https://huggingface.co/blog/nvidia/gemma4#webcamWebcam
v4l2-ctl --list-devices
Usually index0(i.e./dev/video0).
https://huggingface.co/blog/nvidia/gemma4#quick-testQuick test
export MIC_DEVICE="plughw:3,0"
export SPK_DEVICE="alsa_output.usb-Generic_USB2.0_Device_20130100ph0-00.analog-stereo"
arecord -D "$MIC_DEVICE" -f S16_LE -r 16000 -c 1 -d 3 /tmp/test.wav
paplay --device="$SPK_DEVICE" /tmp/test.wav
If you hear yourself, you’re set.
https://huggingface.co/blog/nvidia/gemma4#step-6-run-the-demoStep 6: Run the demo
Make sure the server from Step 4 is running, then:
source .venv/bin/activate
export MIC_DEVICE="plughw:3,0"
export SPK_DEVICE="alsa_output.usb-Generic_USB2.0_Device_20130100ph0-00.analog-stereo"
export WEBCAM=0
export VOICE="af_jessica"
python3 Gemma4_vla.py
On first launch, the script downloads Parakeet STT, Kokoro TTS, and generates voice prompt WAVs. Takes a minute, then you’re live.
- SPACE→ start recording
- speak your question
- SPACE→ stop recording
There’s also a text-only mode if you want to skip audio setup and test the LLM path directly:
python3 Gemma4_vla.py --text
https://huggingface.co/blog/nvidia/gemma4#changing-the-voiceChanging the voice
Kokoro ships with many voices. Switch with:
export VOICE="am_puck"
python3 Gemma4_vla.py
Some good ones:af\_jessica,af\_nova,am\_puck,bf\_emma,am\_onyx.
https://huggingface.co/blog/nvidia/gemma4#how-it-worksHow it works
The script exposes exactly one tool to Gemma 4:
{
"name": "look_and_answer",
"description": "Take a photo with the webcam and analyze what is visible."
}
When you ask a question:
- Your speech is transcribed locally (Parakeet STT)
- Gemma gets the text plus the tool definition
- If the question needs vision, she calls
look\_and\_answer, the script grabs a webcam frame and sends it back - Gemma answers, and Kokoro speaks it out loud
There’s no keyword matching. The model decides when it needs to see. That’s the VLA part.
The\-\-jinjaflag onllama\-serveris what enables this, it activates Gemma’s native tool-calling support.
https://huggingface.co/blog/nvidia/gemma4#troubleshootingTroubleshooting
Server runs out of memory, do the cleanup from Step 3 again. Close everything. This model fits on 8 GB, but you have to be tidy.
No sound, checkpactl list short sinksand make sureSPK\_DEVICEmatches a real sink.
Mic records silence, double-check witharecord \-l, then test recording manually.
First run is slow, normal. It’s downloading models and generating voice prompts. Second run is fast.
https://huggingface.co/blog/nvidia/gemma4#environment-variablesEnvironment variables
VariableDefaultDescriptionLLAMA\_URL``http://127\.0\.0\.1:8080/v1/chat/completionsllama-server endpointMIC\_DEVICE``plughw:3,0ALSA capture deviceSPK\_DEVICE``alsa\_output\.usb\-\.\.\.analog\-stereoPulseAudio sink for playbackWEBCAM``0Webcam index (/dev/videoN)VOICE``af\_jessicaKokoro TTS voice
https://huggingface.co/blog/nvidia/gemma4#bonus-just-want-to-try-gemma-4-in-text-modeBonus: just want to try Gemma 4 in text mode?
If you don’t care about the full VLA demo and just want to poke at Gemma 4 on your Jetson without building anything, there’s a ready-to-go Docker image from Jetson AI Lab withllama\.cpppre-compiled for Orin:
sudo docker run -it --rm --pull always \
--runtime=nvidia --network host \
-v $HOME/.cache/huggingface:/root/.cache/huggingface \
ghcr.io/nvidia-ai-iot/llama_cpp:latest-jetson-orin \
llama-server -hf unsloth/gemma-4-E2B-it-GGUF:Q4_K_S
One line, no compilation, and\-hfpulls the GGUF from Hugging Face on first run. Hithttp://localhost:8080with any OpenAI-compatible client and chat away.
Heads up: this Docker path istext-only. It doesn’t load the vision projector, so it won’t work with the VLA demo above. For the full webcam experience, stick with the native build in Step 4.
Hope you enjoyed this tutorial! If you have any questions or ideas, feel free to reach out. :)
Asier Arranz| NVIDIA
Similar Articles
Welcome Gemma 4: Frontier multimodal intelligence on device
Google DeepMind releases Gemma 4, a frontier multimodal model family available on Hugging Face with Apache 2 licensing, optimized for on-device deployment and supported by various inference libraries.
@googlegemma: Real-time social robotics, from the cloud to your local device. Watch Ian from our DevX team use Gemini Live for a seam…
Google Gemma team demonstrates real-time social robotics using Gemini Live on the Reachy Mini robot, showcasing both cloud and local inference with Gemma 4.
Gemma 4 running fully offline on WebGPU with Transformers.js, controlling Reachy Mini over WebSerial.
Demonstrates running Gemma 4 offline in the browser using WebGPU and Transformers.js to control a Reachy Mini robot via WebSerial.
Gemma 4 E2B running in-browser at 255 tok/s using WebGPU kernels written by Fable 5
Gemma 4 is demonstrated running in-browser via WebGPU at 255 tokens per second, using kernels generated by Fable 5, showcasing efficient on-device inference.
Google’s Gemma 4 12B just dropped - here’s how to run it locally on your Mac
Google released Gemma 4 12B, an Apache 2.0 open-source multimodal model supporting text, vision, and audio with a 256K context window. The article provides a guide for running it locally on Macs using Ollama, LM Studio, or llama.cpp.