gpu-inference

Tag

Cards List
#gpu-inference

Ornith-397B running at Q4 on a single RTX PRO 6000 Blackwell 96GB - 2,354 tok/s prefill, ~20–24 tok/s decode

Reddit r/LocalLLaMA ↗ · 2026-07-27

Krasis, a MoE-focused runtime, enables running the 397B-parameter Ornith model on a single RTX PRO 6000 Blackwell 96GB GPU with ~20-24 tok/s decode by dynamically managing expert residency in VRAM.

0 favorites 0 likes
#gpu-inference

@sudoingX: every day someone asks how i'm running bonsai 27b on hardware that shouldn't handle it. so here's the whole thing in on…

X AI KOLs Timeline ↗ · 2026-07-20 Cached

A detailed guide on running the 27B Bonsai model on hardware with only 8GB VRAM using a 1-bit quantized version and the PrismML fork of llama.cpp, including exact server commands and configuration.

0 favorites 0 likes
#gpu-inference

Accelerating GPU Inference of Large Language Models with Moderately Unstructured Sparse Weight Matrices

arXiv cs.LG ↗ · 2026-07-13 Cached

This paper proposes an efficient GPU inference method for LLMs with moderate unstructured sparsity, introducing a three-layer matrix storage format and a SpMM kernel that jointly utilizes sparse tensor cores and CUDA cores, achieving up to 1.64× kernel-level speedup over SpInfer and up to 1.41× end-to-end speedup over FlashLLM.

0 favorites 0 likes
#gpu-inference

Benchmark of the new unsloth/Qwen3.6-27B-NVFP4 on 4x 5060 ti's with P2P and PP=4 at 1,4,8,12, and 16 concurrency.

Reddit r/LocalLLaMA ↗ · 2026-07-11

Benchmark results of the unsloth/Qwen3.6-27B-NVFP4 model running on 4x RTX 5060 Ti GPUs with peer-to-peer and pipeline parallelism at various concurrency levels.

0 favorites 0 likes
#gpu-inference

@Tono_Ken3: On GPU, driving Qwen3.6-27b (100 TPS) and on CPU, Hy3-299B (25 TPS) simultaneously The dawn of a new era in local LLM i…

X AI KOLs Timeline ↗ · 2026-07-08

Demonstration of running Qwen3.6-27b on GPU at 100 TPS and Hy3-299B on CPU at 25 TPS simultaneously, marking a milestone in local LLM inference.

0 favorites 0 likes
#gpu-inference

6x P40 running Minimax M2.7_Q3_XL

Reddit r/LocalLLaMA ↗ · 2026-07-02

A detailed home lab setup with 6x P40 GPUs running a quantized MiniMax M2.7 model, including hardware specs, benchmark results, and optimal configuration using llama.cpp.

0 favorites 0 likes
#gpu-inference

Qwen3.6-35B-A3B APEX on a Single RTX 3090 - Getting the Most Out of It

Reddit r/LocalLLaMA ↗ · 2026-06-22

A detailed guide on running the Qwen3.6-35B-A3B APEX model on an RTX 3090, comparing two llama.cpp forks and quantization methods for optimal speed and quality.

0 favorites 0 likes
#gpu-inference

@Akashi203: i open-sourced automegakernel -- compiles any huggingface model into a single persistent megakernel batch-1 decode is b…

X AI KOLs Timeline ↗ · 2026-06-17 Cached

AutoMegaKernel is an open-source agent harness that compiles any HuggingFace model into a single persistent megakernel, fusing the entire forward pass into one GPU launch to reduce overhead. It achieves up to 1.33x speedup over CUDA-graphed cuBLAS on inference-class GPUs like L4 and L40S, while proving schedules deadlock- and race-free.

0 favorites 0 likes
#gpu-inference

@PyTorch: ExecuTorch now has an MLX delegate that runs PyTorch models on Apple Silicon GPUs. It supports LLMs, speech-to-text, an…

X AI KOLs Following ↗ · 2026-05-18 Cached

ExecuTorch now has an MLX delegate that enables GPU-accelerated inference for PyTorch models on Apple Silicon Macs, supporting LLMs, speech-to-text, and MoE models with quantization via TorchAO.

0 favorites 0 likes
#gpu-inference

qwen3.6 just stops

Reddit r/LocalLLaMA ↗ · 2026-05-13

A user reports an issue where the Qwen 3.6 model stops mid-task when served via vLLM with specific Docker and speculative decoding configurations.

0 favorites 0 likes
#gpu-inference

Benchmark Qwen 3.6 27B MTP on 2x3090 NVLINK

Reddit r/LocalLLaMA ↗ · 2026-05-08

A benchmark analysis of Qwen 3.6 27B MTP on 4x RTX 3090 GPUs, demonstrating that using NVLink for tensor parallelism yields significant throughput improvements (up to +53%) over PCIe configurations.

0 favorites 0 likes
#gpu-inference

@anyscalecompute: In this session, you'll learn: - Build and scale data pipelines with Ray - What is video data curation - Stream large d…

X AI KOLs Following ↗ · 2026-05-07 Cached

Anyscale is hosting a hands-on virtual lab session teaching developers how to build and scale data pipelines with Ray, covering video data curation, distributed GPU inference, and CPU/GPU streaming pipelines.

0 favorites 0 likes
#gpu-inference

@iotcoi: Qwen3.6-27B-FP8 + Dflash + DDTree, 256k context, 10 agents ~200 tokens/sec max decode 136t/s average on a single tiny G…

X AI KOLs Timeline ↗ · 2026-04-22 Cached

Quantized 27B Qwen3.6 model achieves 200 tok/s peak (136 avg) with 256k context and 10 agents on a single 49W GB10 GPU using Dflash+DDTree optimizations.

0 favorites 0 likes
#gpu-inference

@ProTekkFZS: Q4_K_M 3.6 35B at 768k with yarn on my 3090 has been a joy, I can't lie. Using the llama.cpp fork from @no_stp_on_snek …

X AI KOLs Following ↗ · 2026-04-20 Cached

User reports successfully running a 35B-parameter mixture-of-experts model at 768K context length using Q4_K_M quantization and YaRN on an RTX 3090 via a llama.cpp fork, offloading only 8 experts to CPU while maintaining acceptable performance.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback