@dee_hw: What’s it like to run a frontier model locally at 300 tok/s? Autonomous Computer × DeepSeek V4 Flash is now my daily dr…
Summary
A tweet and product page promoting Autonomous Computer 2, a local AI workstation with dual RTX 5090 GPUs, designed to run frontier models like DeepSeek V4 Flash at high speed with privacy and no per-token costs.
View Cached Full Text
Cached at: 08/12/26, 12:26 PM
What’s it like to run a frontier model locally at 300 tok/s?
Autonomous Computer × DeepSeek V4 Flash is now my daily driver. I’m more efficient. I think faster. I ship more. It feels like a cheat code.
https://t.co/4iRYRPoaVj
I don’t miss Opus or GPT at all. Local AI is here. https://t.co/W0GGvdYNpX
Autonomous
Source: https://www.autonomous.ai/computer-2
Own your intelligence.
Never revoked.
Your models don’t get deprecated, rate-limited, or pulled. The weights sit on your disk.
Private by physics.
Prompts never leave the building. Nobody trains on your work. Not by policy, by physics.
No per-token bill.
Own the machine and agents can run around the clock. You pay for electricity, not tokens.
Yours to keep.
Fine-tune on what only you know. The model becomes an asset you own, and it sells with the company.
GPUs
Two cards, tuned and burn-tested. We installRTX 5090 or RTX PRO 6000 Blackwell— installed, driver-loaded, and benchmarked as one system before it ships. Plug it in and run.
3,584GB/s
Memory bandwidth
Chassis
**Milled from solid aluminum.**Not a bent-steel box — a billet, milled inside and out until only the machine is left. It moves heat like a heatsink; the triangular cutouts keep it stiff without dead weight. 12.5 × 12.5 × 16 inch, anodized black.


Riser
**A full ×16 to every card.**Each GPU reaches the board over its own server-grade PCIe 5.0 riser — shielded, custom-cut, no dropped lanes. Every power run is cut to length and sleeved by hand.
Motherboard
**ASUS Pro WS W790E-SAGE SE.**The hardest choice in a multi-GPU build, solved with a workstation board: every card gets a full ×16, and the BIOS ships pre-tuned so every card links at full width.
CPU
**Intel Xeon w5-3423.**The lanes are the point: 112 lanes of PCIe 5.0 means both cards, the NVMe, and the risers never fight for bandwidth.

Power
**2000W that doesn’t flinch.**Under full load the box pulls up to 1600W — and cheap power is where multi-GPU builds die. We spec 2000W with native 12V-2x6, feeding clean power to every card. The result is boring: it just runs.
Cooling
**Air moved on purpose.**Six fans on a single controller, mapped to the airflow path before a panel was cut — not bolted on after.
Memory & Storage
**ECC memory, fast NVMe.**Error-correcting memory catches the silent bit-flip before it corrupts the job you left running overnight. The NVMe loads a 70B checkpoint in seconds, not minutes. Need more? Configure up to 256GB RAM and 8TB Gen5 NVMe.
4 users76.8Total tok/s840TTFT (s)0.64 s
8 users72.1Total tok/s1,488TTFT (s)1.15 s
16 users64Total tok/s2,390TTFT (s)2.27 s
24 users52Total tok/s3,004TTFT (s)2.41 s
Qwen3.6-27B, per user (tok/s)
Grey: Mac Studio M3 Ultra single user, published run. Blue: this box on vLLM FP8, per user while serving 4 and 24 users at once.
Mac Studio M3 Ultra 512GB, 1 user (MLX 4-bit)
Mac Studio M3 Ultra, 1 user, best tuned build (MTPLX v2)
2x RTX 5090, EACH of 4 users (this box)
2x RTX 5090, EACH of 24 users (this box)
What Qwen3.6-27B gets done in one hour.### Runs these models
Language6
- Qwen3.6 27B
- Qwen3.6 35B A3B
- Gemma-4 31B
- Gemma-4 26B A4B
- Laguna-S 2.14-bit GGUF
- Laguna-XS 2.1
Image3
- Qwen Image 2512
- Qwen Image Editing 2511
- FLUX.2-dev
Video2
- Wan 2.2
- MiniMax H3
Quantized builds noted per model. What fits depends on context length and KV cache.
GPU
Cards
2x NVIDIA RTX 5090 or 2x NVIDIA RTX PRO 6000 Blackwell
Total VRAM
64 GB (RTX 5090) / 192 GB (RTX PRO 6000)
FP32 compute
210 TFLOPS (RTX 5090) / 252 TFLOPS (RTX PRO 6000)
Interconnect
PCIe 5.0 x16 per card, Gen5 riser cables
Compute
CPU
Intel Xeon w5-3423 - 12 cores, 112 PCIe 5.0 lanes
Motherboard
ASUS Pro WS W790E-SAGE SE - 7x PCIe 5.0 x16
Memory & storage
RAM
128 GB (4x 32GB DDR5-4800 ECC) - configurable 64-256 GB
Storage
2TB NVMe PCIe 4.0 - configurable 1-8TB, Gen5 on 8TB
Power & cooling
Fans
6x Thermalright TL-N12-R9, single controller
Dimensions
Chassis
CNC-machined solid aluminum, anodized black
Software
OS
Ubuntu, NVIDIA drivers pre-installed
AI frameworks
PyTorch, CUDA, vLLM, SGLang, llama.cpp
Warranty
Other
Open-source hardware
The whole machine is open.
Every CAD file, bill of materials, and BIOS setting — free on GitHub. Fork it, change it, build your own, even sell it. This is how the personal computer began: in the open. We’re doing it again, for AI hardware.
autonomous-ai / autonomous-computer
2x-5090/ 4x-5090/ 8x-5090/ 4x-6000/
bom/ step_models/ stl-models/ photos/
README.md setup.md
MIT · 147 forks
Clone it. Build it. Sell it. We just want it built.
Similar Articles
@danveloper: https://x.com/danveloper/status/2064387956387758206
A developer ran DeepSeek-V4-Flash on a Raspberry Pi 5 by streaming model weights from an NVMe SSD, achieving 1.3 tokens/second at 8 watts, demonstrating the feasibility of frontier-adjacent open-weight models on low-cost, offline hardware.
Running DeepSeek-V4 locally with 4x legacy RTX 2080 Ti ($2k budget setup). Custom Turing kernels, W8A8 quantization, and 255 prefill tok/s!
A developer successfully runs DeepSeek-V4-Flash (284B total, 13B active) locally on four RTX 2080 Ti GPUs with a $2,500 budget, achieving 255 prefill tokens/s using custom Turing CUDA kernels, W8A8 quantization, and heterogeneous inference. The implementation is open-sourced.
@ciruai: Testing DeepSeek v4 Flash on the AMD Ryzen AI Max+ 395 Strix Halo with 128GB RAM. Getting ~15 TPS over a decently long …
Testing DeepSeek v4 Flash on the AMD Ryzen AI Max+ 395 with 128GB RAM achieves ~15 TPS for a 284B MoE model (13B active) locally, costing $3,000 versus $25,000+ for a datacenter setup, highlighting the feasibility of running large models on consumer hardware.
I CANNOT believe I've got DeepSeek-V4-Flash-0731, a frontier model, running on my home PC. Insane!
A user expresses astonishment at running DeepSeek-V4-Flash-0731, a frontier model, on a mid-range Windows PC with 24GB VRAM via quantization, highlighting rapid progress in local AI.
@dee_hw: Qwen3.8-27B is coming. We open source our 2x RTX 5090 build, so you can host it locally: on-prem, private, no rate limi…
Autonomous AI open-sources a Personal AI Computer build powered by 2x or 4x NVIDIA RTX 5090s, enabling fully local, private AI hosting without API costs or rate limits.