gpu-inference

Tag

Cards List
#gpu-inference

tp=6 can work on vLLM, with padding

Reddit r/LocalLLaMA ↗ · 13h ago

A user shares a workaround enabling vLLM tensor parallelism with tp=6 on six GPUs by padding model architecture dimensions with zeros until divisible, achieving higher KV cache utilization for Qwen 27B inference on Radeon 7900 XTX GPUs.

0 favorites 0 likes
#gpu-inference

@coldniko: Run Qwen3.8-Flash-next on 8GB+ AMD GPU RX 7900 XTX runs the model at a sustained 52-60 output tokens per second + 1250 …

X AI KOLs Timeline ↗ · yesterday Cached

Strata is an open-source tool that runs Qwen3.8-Flash-Next, a 125B-parameter model, on consumer gaming PCs with quantization, delivering 52-60 tokens/s on an AMD RX 7900 XTX and 60-95 tokens/s on an RTX 5070, with KV cache optimizations like k8v4 for an extra 10% speedup.

0 favorites 0 likes
#gpu-inference

@DavidFSWD: THE MOST Amzing thing in Local AI this week is this optimized Qwen 3.8 model: ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF…

X AI KOLs Timeline ↗ · 4d ago Cached

This tweet highlights an optimized version of the Qwen 3.8 model that can run efficiently on local hardware with under 16GB VRAM, achieving 30-80 tokens per second on older GPU setups.

0 favorites 0 likes
#gpu-inference

To the dozens of 3x 3090 Local LLM people - I found our current best fit

Reddit r/LocalLLaMA ↗ · 2026-09-20

The author finds that running Qwen 3.8 Next Flash on Exllama3 at 3.05 bpw on 3x 3090 GPUs delivers exceptional performance and quality for local LLM usage, outperforming other quantizations.

0 favorites 0 likes
#gpu-inference

Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling

arXiv cs.LG ↗ · 2026-09-18 Cached

This paper explores the impact of different candidate-generation schedules on the energy consumption and performance of large language models during test-time scaling, demonstrating that larger batch sizes reduce energy use and latency.

0 favorites 0 likes
#gpu-inference

@BenjaminDEKR: I'm testing this on a 5070 Ti (16gb vram) now and it's really impressive so far

X AI KOLs Following ↗ · 2026-09-18 Cached

A user is testing the newly announced Ternary Bonsai 2 27B AI model, a smaller and quantized version of Qwen3.8 27B, on an NVIDIA 5070 Ti GPU and finds its performance impressive.

0 favorites 0 likes
#gpu-inference

It started as a GitHub profile generator. Then it accidentally traded stocks. Now it’s an AI agentic platform

Reddit r/AI_Agents ↗ · 2026-09-13

Magine evolved from a simple GitHub profile generator into an AI agentic platform featuring sight-driven agents and a distributed GPU mesh, with growth metrics and technical innovations in agent orchestration.

0 favorites 0 likes
#gpu-inference

Voice conversations between Gemma4 12B and E2B on GPU and Jetson Orin

Reddit r/LocalLLaMA ↗ · 2026-09-08

The article demonstrates voice conversations between Gemma 4 models on GPU and Jetson Orin hardware, using the open-source Cortexist Little Gemma engine for efficient inference with lip sync and gestures.

0 favorites 0 likes
#gpu-inference

(NInfer Fork) I wanted to have a 1M context Qwen-3.8 27B, tp2, dual 5090s

Reddit r/LocalLLaMA ↗ · 2026-08-29

A developer forked NInfer, a C++20/CUDA inference engine, to add tensor-parallelism and YaRN rope scaling, enabling Qwen3.8-27B to run with a 1M token context on dual 5090 GPUs and outperforming vLLM in specific decode scenarios.

0 favorites 0 likes
#gpu-inference

yall are sleeping on qwen 3.8 27b q2 + q2 dflash + q5 kv

Reddit r/LocalLLaMA ↗ · 2026-08-28

A user shares their experience running a quantized Qwen 3.8 27B model using QAT Q2 and Q5 KV, achieving high performance on a 12GB GPU with up to 200K token context, surpassing models like Sonnet 4.6.

0 favorites 0 likes
#gpu-inference

Qwen 3.8 27B is faster than expected

Reddit r/LocalLLaMA ↗ · 2026-08-18

A user reports that the Qwen 3.8 27B model achieves 50-60 tokens per second on dual 5060 TI cards, showing unexpected speed improvements over previous versions like Qwen 3.6.

0 favorites 0 likes
#gpu-inference

Qwen3.8 vs Qwen3.6 vs Gemma 4 on a 24GB GPU (10 minute read)

TLDR AI ↗ · 2026-08-18 Cached

This article benchmarks and compares the performance of Qwen3.8-27B, Qwen3.6-27B, and Gemma 4 31B on a 24GB GPU, recommending Qwen3.8-27B as the best default for most users due to superior coding and reasoning capabilities.

0 favorites 0 likes
#gpu-inference

Qwen3.8-27B lands next to DeepSeek V4 and GPT-5.6 Luna Max on the Artificial Analysis Benchmark. You can now run a near frontier model with just a RTX 3090.

Reddit r/singularity ↗ · 2026-08-17 Cached

Qwen3.8-27B is a new AI model that achieves competitive performance with frontier models like DeepSeek V4 and GPT-5.6 Luna Max on Artificial Analysis benchmarks, and it can be run locally on an RTX 3090 GPU.

0 favorites 0 likes
#gpu-inference

100$ worth of gpu runs qwen 3.8 27b at 7.39 t/s

Reddit r/LocalLLaMA ↗ · 2026-08-17

A user demonstrates running the Qwen 27b AI model quantized to Q3_K_M on two RX 580 GPUs, achieving 7.39 tokens per second using old DDR3 hardware for under $100.

0 favorites 0 likes
#gpu-inference

How accurate do you think this is? Qwen3.5 9B vs GPT-4o

Reddit r/LocalLLaMA ↗ · 2026-08-17

The post questions whether Qwen 3.5 9B running on low-resource systems can outperform GPT-4o, highlighting its small size of less than 7 GB and claimed superior performance.

0 favorites 0 likes
#gpu-inference

Genie-style playable world model running 720p at 16 FPS on a single 5090 in 19GB VRAM

Reddit r/LocalLLaMA ↗ · 2026-08-16 Cached

ABot-World-0 is a real-time interactive world model that achieves 720p video generation at 16 FPS on a single NVIDIA RTX 5090 GPU with 19GB VRAM, enabling infinite action-conditioned world rollout for simulation and AI research.

0 favorites 0 likes
#gpu-inference

I tested the CMP170HX

Reddit r/LocalLLaMA ↗ · 2026-08-11

A hands-on benchmark of Nvidia CMP170HX mining cards repurposed as 64GB VRAM AI inference accelerators, showing they can run large local LLMs like DeepSeek V4-Flash and gpt-oss-120B at useful speeds, with caveats around Ampere-class throughput and PCIe Gen2 x4 connectivity.

0 favorites 0 likes
#gpu-inference

@TheAhmadOsman: Dense models like Qwen 3.8 27B are a TERRIBLE experience on unified-memory systems like DGX Spark btw DGX Sparks are be…

X AI KOLs Following ↗ · 2026-08-08 Cached

Ahmad Osman argues that dense models like Qwen 27B perform poorly on unified-memory systems such as NVIDIA DGX Spark, suggesting MoE models are a better fit; he claims discrete GPUs like the RTX PRO 6000 deliver far better performance for agentic workloads.

0 favorites 0 likes
#gpu-inference

Topology-Aware Data Movement for Disaggregated GPU Inference

arXiv cs.LG ↗ · 2026-08-03 Cached

This paper presents a topology-aware data movement orchestrator for disaggregated LLM inference, which dynamically selects optimal transport based on interconnect hierarchy and overlaps KV cache transfer with computation, achieving 3-18x transfer latency reduction over uniform RDMA.

0 favorites 0 likes
#gpu-inference

Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution

Hacker News Top ↗ · 2026-07-29 Cached

An updated benchmark shows self-hosting Kimi K3 on 8×B300 nodes achieves 86.4% task resolution at roughly 20% higher hardware cost compared to GLM-5.2 on B200 nodes, though with lower throughput.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback