Tag
The user built a low-cost setup using five ex-mining BC-250 boards to run the Qwen3-Coder-Next AI model, achieving around 40 tokens per second at 30k context with plans to expand.
The article describes a project at an art residence where a local LLM processes live audio and visual inputs from a party to generate dynamic projections on a wall, exploring how people interact with the AI-driven environment.
A user shares detailed benchmarking data and personal insights on running AI models locally with varying GPU power limits, evaluating models like gemma4 and qwen3.5 on a modest hardware setup.
A user optimized a 4x3060ti GPU rig for AI inference using tensor parallelism with Exl3 and vllm, achieving up to 120 tokens per second with large context windows.
A user shares their local AI setup using two BC-250 ex-mining APUs to run the Qwen3.6-35B-A3B model with llama.cpp, achieving 60 tok/s and 64k context for under $300.
The author finds that running Qwen 3.8 Next Flash on Exllama3 at 3.05 bpw on 3x 3090 GPUs delivers exceptional performance and quality for local LLM usage, outperforming other quantizations.
A person shares that their closet is turning into a micro data center, requiring watercooling and power generators.
A user revived an old ThinkPad T480 by installing Omarchy, OpenCode, and Local AI Grid to run DeepSeek V4F locally at 300 tok/s for coding purposes.
The article details a validated hardware configuration using 16 RTX 5060 Ti GPUs with PLX switches to run the Deepseek V4 Flash model, achieving specific performance metrics for context handling and throughput.
This article documents the complete process of the author remotely connecting to the company's DGX Spark server on the weekend to successfully deploy the Ling-3.0-flash model, including selection, deployment, performance testing, and integration with development tools, and shares insights on local deployment as a controllable intermediate state.
The article shares an optimal llama.cpp configuration for running the Qwen 3.8 27B model on 16GB VRAM with 73k context, demonstrating its performance in agentic coding workflows through a real-world software engineering project.
A user outlines a step-by-step plan to achieve 200GB of VRAM by combining multiple NVIDIA GPUs in a custom PC build, addressing purchase, installation, and power management.
A new wave of large language models including GLM 4.5, Qwen 3.5, MiniMax M2.7, Deepseek V4 Flash, Xiaomi MiMo 2.5, StepFun 3.7 Flash, and Tencent Hy3 can now be run locally on a dual DGX Spark setup with 250GB usable memory at 4-bit quantization, costing approximately $7,000–$8,000.
Testing AI models Qwen3, Qwen3.6, and Gemma4 on a CPU setup using the Intel 285hx processor (MS-02).
This post details running GLM 5.2 on a 4xGB10 setup with a 100G switch, achieving ~25 tok/s decode and ~650 tok/s prefill at 330k context. It includes hardware costs, performance benchmarks with Depth Prefill, and notes on model pruning for longer context.
A user describes their fully local AI stack using multiple hardware devices running Chinese models like GLM, Qwen, and Kimi, claiming 87% cost savings compared to frontier models like GPT-5.5 and Opus 4.8, while noting plans to self-host video generation.
Discusses running a Q6 quantized version of the Gemma 4 31B model on a dual 9060 XT GPU configuration, likely for local inference.