hardware-setup

Tag

Cards List
#hardware-setup

I’m calling this the Monstrosity. 5 ex mining BC-250 boards Qwen3-Coder-Next Q4 at 40 tok/s

Reddit r/LocalLLaMA ↗ · 2026-09-28

The user built a low-cost setup using five ex-mining BC-250 boards to run the Qwen3-Coder-Next AI model, achieving around 40 tokens per second at 30k context with plans to expand.

0 favorites 0 likes
#hardware-setup

We let a local LLM run the wall at a party, and people started performing for it. Notes from the first test

Reddit r/ArtificialInteligence ↗ · 2026-09-27

The article describes a project at an art residence where a local LLM processes live audio and visual inputs from a party to generate dynamic projections on a wall, exploring how people interact with the AI-driven environment.

0 favorites 0 likes
#hardware-setup

Power Limits, Local AI, and Questionable Uses of My Free Time

Reddit r/LocalLLaMA ↗ · 2026-09-27

A user shares detailed benchmarking data and personal insights on running AI models locally with varying GPU power limits, evaluating models like gemma4 and qwen3.5 on a modest hardware setup.

0 favorites 0 likes
#hardware-setup

Getting stupidly good results on my 4x3060ti setup.

Reddit r/LocalLLaMA ↗ · 2026-09-26

A user optimized a 4x3060ti GPU rig for AI inference using tensor parallelism with Exl3 and vllm, achieving up to 120 tokens per second with large context windows.

0 favorites 0 likes
#hardware-setup

My foray into local ai. Two BC-250 ex mining apus running Qwen3.6-35B-A3B Q4_K_M at 60 tok/s with 64k context

Reddit r/LocalLLaMA ↗ · 2026-09-24

A user shares their local AI setup using two BC-250 ex-mining APUs to run the Qwen3.6-35B-A3B model with llama.cpp, achieving 60 tok/s and 64k context for under $300.

0 favorites 0 likes
#hardware-setup

To the dozens of 3x 3090 Local LLM people - I found our current best fit

Reddit r/LocalLLaMA ↗ · 2026-09-20

The author finds that running Qwen 3.8 Next Flash on Exllama3 at 3.05 bpw on 3x 3090 GPUs delivers exceptional performance and quality for local LLM usage, outperforming other quantizations.

0 favorites 0 likes
#hardware-setup

@nikitabier: My closet is turning into a micro data center. I will soon need watercooling and power generators. Not sure how I ended…

X AI KOLs Timeline ↗ · 2026-09-18

A person shares that their closet is turning into a micro data center, requiring watercooling and power generators.

0 favorites 0 likes
#hardware-setup

@dee_hw: Found this ThinkPad T480 in the garage. I just set it up with: Omarchy: http://omarchy.org OpenCode: http://opencode.ai…

X AI KOLs Timeline ↗ · 2026-08-24 Cached

A user revived an old ThinkPad T480 by installing Omarchy, OpenCode, and Local AI Grid to run DeepSeek V4F locally at 300 tok/s for coding purposes.

0 favorites 0 likes
#hardware-setup

The boring way to run Deepseek V4 Flash-0731 130-150 tks - 16x5060ti 16GB over 2 PLX88096 switches

Reddit r/LocalLLaMA ↗ · 2026-08-20

The article details a validated hardware configuration using 16 RTX 5060 Ti GPUs with PLX switches to run the Deepseek V4 Flash model, achieving specific performance metrics for context handling and throughput.

0 favorites 0 likes
#hardware-setup

@xiaomovps: After a company starts using AI, they quickly hit several hard problems: whether data can be externalized, whether costs can be controlled, and whether to build internal models themselves. This article documents a very real weekend operation—remotely connecting to the company's DGX Spark and hands-on running Ling-3.0-flash. Not stopping at 'can it run', but…

X AI KOLs Timeline ↗ · 2026-08-18 Cached

This article documents the complete process of the author remotely connecting to the company's DGX Spark server on the weekend to successfully deploy the Ling-3.0-flash model, including selection, deployment, performance testing, and integration with development tools, and shares insights on local deployment as a controllable intermediate state.

0 favorites 0 likes
#hardware-setup

After pushing 1M+ tokens through Qwen 3.8 27B, here is my optimal llama.cpp config for 16GB VRAM (73k Context, Agentic Coding)

Reddit r/LocalLLaMA ↗ · 2026-08-17

The article shares an optimal llama.cpp configuration for running the Qwen 3.8 27B model on 16GB VRAM with 73k context, demonstrating its performance in agentic coding workflows through a real-world software engineering project.

0 favorites 0 likes
#hardware-setup

The dream is to reach 200GB VRAM

Reddit r/LocalLLaMA ↗ · 2026-08-16

A user outlines a step-by-step plan to achieve 200GB of VRAM by combining multiple NVIDIA GPUs in a custom PC build, addressing purchase, installation, and power management.

0 favorites 0 likes
#hardware-setup

New wave of miniboss models you can run on dual DGX Spark

Reddit r/LocalLLaMA ↗ · 2026-07-15

A new wave of large language models including GLM 4.5, Qwen 3.5, MiniMax M2.7, Deepseek V4 Flash, Xiaomi MiMo 2.5, StepFun 3.7 Flash, and Tencent Hy3 can now be run locally on a dual DGX Spark setup with 250GB usable memory at 4-bit quantization, costing approximately $7,000–$8,000.

0 favorites 0 likes
#hardware-setup

First attempts at a CPU setup - MS-02 Intel 285hx, trying Qwen3, Qwen3.6 and Gemma4

Reddit r/LocalLLaMA ↗ · 2026-07-12

Testing AI models Qwen3, Qwen3.6, and Gemma4 on a CPU setup using the Intel 285hx processor (MS-02).

0 favorites 0 likes
#hardware-setup

Running GLM 5.2 on 4xGB10 with a 100G Switch, 330k ctx, ~25 t/s tg, ~650 t/s pp

Reddit r/LocalLLaMA ↗ · 2026-07-08

This post details running GLM 5.2 on a 4xGB10 setup with a 100G switch, achieving ~25 tok/s decode and ~650 tok/s prefill at 330k context. It includes hardware costs, performance benchmarks with Depth Prefill, and notes on model pruning for longer context.

0 favorites 0 likes
#hardware-setup

@DeRonin_: My current local AI setup: - 2x DGX Spark linked (256gb) > GLM 5.2 @ 2bit, reasoning + agent loops - Mac Studio M3 Ultr…

X AI KOLs Following ↗ · 2026-06-30 Cached

A user describes their fully local AI stack using multiple hardware devices running Chinese models like GLM, Qwen, and Kimi, claiming 87% cost savings compared to frontier models like GPT-5.5 and Opus 4.8, while noting plans to self-host video generation.

0 favorites 0 likes
#hardware-setup

Gemma 4 31B Q6 on Dual 9060 XT

Reddit r/LocalLLaMA ↗ · 2026-06-22

Discusses running a Q6 quantized version of the Gemma 4 31B model on a dual 9060 XT GPU configuration, likely for local inference.

0 favorites 0 likes
← Back to home

Submit Feedback