@dee_hw: What’s it like to run a frontier model locally at 300 tok/s? Autonomous Computer × DeepSeek V4 Flash is now my daily dr…

X AI KOLs Timeline Products

Summary

A tweet and product page promoting Autonomous Computer 2, a local AI workstation with dual RTX 5090 GPUs, designed to run frontier models like DeepSeek V4 Flash at high speed with privacy and no per-token costs.

What’s it like to run a frontier model locally at 300 tok/s? Autonomous Computer × DeepSeek V4 Flash is now my daily driver. I’m more efficient. I think faster. I ship more. It feels like a cheat code. https://t.co/4iRYRPoaVj I don't miss Opus or GPT at all. Local AI is here. https://t.co/W0GGvdYNpX
Original Article
View Cached Full Text

Cached at: 08/12/26, 12:26 PM

What’s it like to run a frontier model locally at 300 tok/s?

Autonomous Computer × DeepSeek V4 Flash is now my daily driver. I’m more efficient. I think faster. I ship more. It feels like a cheat code.

https://t.co/4iRYRPoaVj

I don’t miss Opus or GPT at all. Local AI is here. https://t.co/W0GGvdYNpX


Autonomous

Source: https://www.autonomous.ai/computer-2

Own your intelligence.

Never revoked.

Your models don’t get deprecated, rate-limited, or pulled. The weights sit on your disk.

Private by physics.

Prompts never leave the building. Nobody trains on your work. Not by policy, by physics.

No per-token bill.

Own the machine and agents can run around the clock. You pay for electricity, not tokens.

Yours to keep.

Fine-tune on what only you know. The model becomes an asset you own, and it sells with the company.

GPUs

Two cards, tuned and burn-tested. We installRTX 5090 or RTX PRO 6000 Blackwell— installed, driver-loaded, and benchmarked as one system before it ships. Plug it in and run.

3,584GB/s

Memory bandwidth

Chassis

**Milled from solid aluminum.**Not a bent-steel box — a billet, milled inside and out until only the machine is left. It moves heat like a heatsink; the triangular cutouts keep it stiff without dead weight. 12.5 × 12.5 × 16 inch, anodized black.

Machined detail

Machined detail

Riser

**A full ×16 to every card.**Each GPU reaches the board over its own server-grade PCIe 5.0 riser — shielded, custom-cut, no dropped lanes. Every power run is cut to length and sleeved by hand.

Motherboard

**ASUS Pro WS W790E-SAGE SE.**The hardest choice in a multi-GPU build, solved with a workstation board: every card gets a full ×16, and the BIOS ships pre-tuned so every card links at full width.

CPU

**Intel Xeon w5-3423.**The lanes are the point: 112 lanes of PCIe 5.0 means both cards, the NVMe, and the risers never fight for bandwidth.

Machined detail

Power

**2000W that doesn’t flinch.**Under full load the box pulls up to 1600W — and cheap power is where multi-GPU builds die. We spec 2000W with native 12V-2x6, feeding clean power to every card. The result is boring: it just runs.

Cooling

**Air moved on purpose.**Six fans on a single controller, mapped to the airflow path before a panel was cut — not bolted on after.

Memory & Storage

**ECC memory, fast NVMe.**Error-correcting memory catches the silent bit-flip before it corrupts the job you left running overnight. The NVMe loads a 70B checkpoint in seconds, not minutes. Need more? Configure up to 256GB RAM and 8TB Gen5 NVMe.

4 users76.8Total tok/s840TTFT (s)0.64 s

8 users72.1Total tok/s1,488TTFT (s)1.15 s

16 users64Total tok/s2,390TTFT (s)2.27 s

24 users52Total tok/s3,004TTFT (s)2.41 s

Qwen3.6-27B, per user (tok/s)

Grey: Mac Studio M3 Ultra single user, published run. Blue: this box on vLLM FP8, per user while serving 4 and 24 users at once.

Mac Studio M3 Ultra 512GB, 1 user (MLX 4-bit)

Mac Studio M3 Ultra, 1 user, best tuned build (MTPLX v2)

2x RTX 5090, EACH of 4 users (this box)

2x RTX 5090, EACH of 24 users (this box)

What Qwen3.6-27B gets done in one hour.### Runs these models

Language6

  • Qwen3.6 27B
  • Qwen3.6 35B A3B
  • Gemma-4 31B
  • Gemma-4 26B A4B
  • Laguna-S 2.14-bit GGUF
  • Laguna-XS 2.1

Image3

  • Qwen Image 2512
  • Qwen Image Editing 2511
  • FLUX.2-dev

Video2

  • Wan 2.2
  • MiniMax H3

Quantized builds noted per model. What fits depends on context length and KV cache.

GPU

Cards

2x NVIDIA RTX 5090 or 2x NVIDIA RTX PRO 6000 Blackwell

Total VRAM

64 GB (RTX 5090) / 192 GB (RTX PRO 6000)

FP32 compute

210 TFLOPS (RTX 5090) / 252 TFLOPS (RTX PRO 6000)

Interconnect

PCIe 5.0 x16 per card, Gen5 riser cables

Compute

CPU

Intel Xeon w5-3423 - 12 cores, 112 PCIe 5.0 lanes

Motherboard

ASUS Pro WS W790E-SAGE SE - 7x PCIe 5.0 x16

Memory & storage

RAM

128 GB (4x 32GB DDR5-4800 ECC) - configurable 64-256 GB

Storage

2TB NVMe PCIe 4.0 - configurable 1-8TB, Gen5 on 8TB

Power & cooling

Fans

6x Thermalright TL-N12-R9, single controller

Dimensions

Chassis

CNC-machined solid aluminum, anodized black

Software

OS

Ubuntu, NVIDIA drivers pre-installed

AI frameworks

PyTorch, CUDA, vLLM, SGLang, llama.cpp

Warranty

Other

Open-source hardware

The whole machine is open.

Every CAD file, bill of materials, and BIOS setting — free on GitHub. Fork it, change it, build your own, even sell it. This is how the personal computer began: in the open. We’re doing it again, for AI hardware.

GitHub1.2k

autonomous-ai / autonomous-computer

2x-5090/ 4x-5090/ 8x-5090/ 4x-6000/

bom/ step_models/ stl-models/ photos/

README.md setup.md

MIT · 147 forks

Clone it. Build it. Sell it. We just want it built.

Similar Articles

@danveloper: https://x.com/danveloper/status/2064387956387758206

X AI KOLs Timeline

A developer ran DeepSeek-V4-Flash on a Raspberry Pi 5 by streaming model weights from an NVMe SSD, achieving 1.3 tokens/second at 8 watts, demonstrating the feasibility of frontier-adjacent open-weight models on low-cost, offline hardware.