Dual rtx 3090 build

Reddit r/LocalLLaMA News

Summary

A user shares their dual RTX 3090 build for local LLM inference and seeks advice on tool stacks for agentic work and RAG pipelines.

Joining this community sparked a new hobby and interest in software engineering that I had lost. So I made this dual rtx 3090 build mostly for inference , I know I won’t be replacing chatgpt anytime soon but what tool stack would help it be usable in a work environment ? Must MCP servers or custom tools/scripts ? Currently using VScode preview with qwen3.6 27b and an nginx server, Im mostly interested in agentic work with usable context or at least a better knowledge of code base ( RAG pipeline?) Been already such a helpful community , hopefully local llms continue to grow because I fear cloud will become unaffordable at a consumer level
Original Article

Similar Articles

NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B on 2x3090s

Reddit r/LocalLLaMA

A detailed guide on running the quantized NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B model on two RTX 3090s using vLLM with full 262K context, achieving high inference speeds without CPU offloading.