Jetson Orin NX Build for Hermes Agent + Benchmarking
Summary
A detailed build and benchmarking of a Jetson Orin NX system for running Hermes Agent, achieving 14.65 tok/s at 8k context and 10.21 tok/s at 60k context with Gemma 4 26B quantized model.
Similar Articles
@mr_r0b0t: hermesbench v0.1 — a benchmark purpose-built for Hermes Agent tool-calling Runs real Hermes Agent subprocess in an isol…
hermesbench v0.1 is a new benchmark for evaluating local models on Hermes Agent tool-calling, featuring real agent harness, deterministic verifiers, and hardware telemetry. It includes 43 tasks across 11 categories and is designed to measure how well models use the Hermes Agent.
@steipete: Looks like our focus on performance paid off.
A comparison shows Hermes Agent outperforms OpenClaw in token processing time and code generation using a local Qwen 35B model on a MacBook Pro M5 Max.
@analogalok: I just got Gemma 4 26B A4B MoE model running fully locally with Hermes agent on an 8GB RTX 4060 and it's now backtestin…
A developer demonstrates running Gemma 4 26B MoE model locally on an 8GB RTX 4060 with Hermes agent to fully automate backtesting of trading strategies, highlighting the growing capability of local LLMs as autonomous agents.
PrismML Bonsai 27B is surprisingly usable on the Jetson Orin Nano 8GB
PrismML's Bonsai 27B model runs on the Jetson Orin Nano 8GB with 4.31 tokens/s and 27 t/s prompt processing, using 6.2GB RAM and about 25W power. It indicates surprisingly usable edge AI performance.
Got a 27B model running locally on a Jetson Orin NX 16GB (1-bit). still kind of amazed it works
User reports successfully running a 27B parameter model quantized to 1-bit on a Jetson Orin NX 16GB edge device, expressing amazement at the feasibility.