Are you ready for Le Chaton FAT or still wasting money on GPUs?
Summary
The author shares their storage server build optimized for local AI inference, anticipating a rumored 26T-a3b model called "Le Chaton FAT" and using high-capacity NVMe drives with ZFS for model storage.
Similar Articles
3k$ 128GB VRAM + 256GB RAM DDR4 Server
A user details building a home inference server with 128GB VRAM and 256GB DDR4 RAM for AI workloads, achieving satisfactory performance with Qwen3.8 models using a VLLM fork after initial setup issues.
Building a budget 32GB → 48GB VRAM home AI server: 2-3x RX 9060 XT 16GB vs RTX 5060 Ti 16GB, AM5 vs used EPYC?
A user seeks advice on building a budget home AI server with 32-48GB VRAM, debating between AMD RX 9060 XT and Nvidia RTX 5060 Ti GPUs, and whether to use AM5 or used EPYC platforms for local LLM inference and large MoE model offloading.
Don't trust frontier models when asking about budget hardware!
The author shares their experience using Tesla P100 GPUs for local AI inference, finding them cost-effective and performant with optimizations via llama.cpp, despite initial advice from frontier models.
V100 4-card AI large model, Tesla 128G server
Announces a server configuration with 4 Nvidia V100 GPUs and 128GB Tesla memory, targeting AI large model workloads.
@andrewchen: finding the main downside with experimenting with local AI models is that you end up buying one GPU, then another, then…
Andrew Chen shares his experience of buying multiple GPUs for local AI experimentation, running Qwen3.6 27B dense at 100 tok/s on a 5090 eGPU, and compares it to Sonnet 4.6.