inference-speed

Tag

Cards List
#inference-speed

10% faster decode with Q4_K MTP draft model with Gemma 4 31b

Reddit r/LocalLLaMA · 3d ago

A user reports that quantising the f16 MTP draft model to Q4_K for Gemma 4 31b gives roughly 10% faster decode (65 to 72 TPS) on dual 3090s compared to the default Q4_0, while Q2_K performs worse.

0 favorites 0 likes
#inference-speed

@YRSM_Simon: 120 t/s ! Good job, @UnslothAI

X AI KOLs Following · 3d ago Cached

Unsloth AI announces DSpark, enabling DeepSeek-V4-Flash GGUF models to run ~1.4–2× faster locally, reaching 120 tokens/s with no accuracy change.

0 favorites 0 likes
#inference-speed

Final optimization: from ~10 tok/s to ~15 tok/s on DeepSeek-V4-Flash-0731 at 128K ctx - 1 RTX 3090

Reddit r/LocalLLaMA · 3d ago

Tests et réglages détaillés pour optimiser DeepSeek-V4-Flash-0731 en GGUF sur une RTX 3090, atteignant ~15 tok/s à 128K de contexte grâce à différentes quantifications et paramètres de chargement.

0 favorites 0 likes
#inference-speed

Would extremely high decode tok/s even be useful?

Reddit r/LocalLLaMA · 2026-07-30

A discussion questioning whether extremely high decode speeds (1k-10k tok/s) for large models like Qwen 3.5 397B or GLM-5.2 would unlock new use cases or be better spent on loading larger models.

0 favorites 0 likes
#inference-speed

@googlegemma: Voice AI without the wait! Thanks to Hugging Face and Cerebras, developers can now use the Gemma 4 31B model as the bra…

X AI KOLs Timeline · 2026-07-20 Cached

Google Gemma announces that developers can now use the Gemma 4 31B model as the brain for voice AI, enabled by Hugging Face and Cerebras for ultra-fast inference, as part of an open-source cascaded speech-to-speech stack.

0 favorites 0 likes
#inference-speed

@rohanpaul_ai: Spatially Speculative Decoding (SSD) sped up autoregressive image models up to 13.28X by predicting image rows in paral…

X AI KOLs Timeline · 2026-07-14 Cached

Spatially Speculative Decoding (SSD) accelerates autoregressive image models by predicting entire rows in parallel using small helper networks, achieving up to 13.28x speedup while maintaining benchmark performance.

0 favorites 0 likes
#inference-speed

@_avichawla: NVIDIA researchers built a new transformer variant. One small change to the layers made: - decoding 1.7x faster - long-…

X AI KOLs Timeline · 2026-07-12 Cached

NVIDIA researchers introduced SparDA, a transformer variant that adds a fourth projection (Forecast) to predict next-layer KV blocks, enabling prefetching from CPU memory and reducing selection cost, achieving 1.7x faster decoding and 6.5 point accuracy gain on long reasoning.

0 favorites 0 likes
#inference-speed

NVIDIA Puzzle-75B-A9B NVFP4 at 132 t/s on 3×3090 — Why is this size category a desert otherwise?

Reddit r/LocalLLaMA · 2026-07-09

NVIDIA's Puzzle-75B-A9B model achieves 132 tokens per second using NVFP4 quantization on three RTX 3090 GPUs, raising discussion about the lack of competition in this model size category.

0 favorites 0 likes
#inference-speed

@superalesha: I sped up deepseek v4 flash by 29x on my 4x3090s !!! No, its not joke. 15 -> 443 t/s. a 23k prompt used to take 25 mins…

X AI KOLs Timeline · 2026-07-07 Cached

A user achieved a 29x speedup for DeepSeek V4 Flash inference on 4x RTX 3090 GPUs by optimizing llama.cpp, reducing a 23k prompt from 25 minutes to 53 seconds.

0 favorites 0 likes
#inference-speed

I tested freshly merged DFlash in llama.cpp on Qwen 3.6 27B Local AI win. 4.44x faster at 36K context. Here are my findings RTX 6000 PRO.

Reddit r/LocalLLaMA · 2026-07-07

A user benchmarks the newly merged DFlash speculative decoding method in llama.cpp on Qwen 3.6 27B, achieving up to 4.44x speedup at 36K context compared to baseline, with detailed leaderboard and quality tests.

0 favorites 0 likes
#inference-speed

@LiorOnAI: You now convert any LLM into a faster one without retraining from scratch. NVIDIA just did this to their 30B model. Her…

X AI KOLs Timeline · 2026-07-01 Cached

NVIDIA proposes a method to convert any LLM into a faster one by splitting it into two copies: one frozen for context, the other trained to generate multiple tokens in parallel, achieving 2.4x speedup with ~99% quality retention using only 8% of training data.

0 favorites 0 likes
#inference-speed

Thinking about grabbing 4x Ascend GX10s

Reddit r/LocalLLaMA · 2026-07-01

A user considers buying four Ascend GX10s to run GLM5.2, citing performance numbers like 400-500 tok/s prompt processing and ~15 tok/s output at 128k context, and plans for future open-source models.

0 favorites 0 likes
#inference-speed

@MiaAI_lab: Nvidia did it again! @NVIDIAAI's Qwen 3.6 27B NVFP4 is faster than Unsloth's Qwen 3.6 27B NVFP4 by a whopping ~41% on D…

X AI KOLs Timeline · 2026-06-30 Cached

Nvidia's optimized Qwen 3.6 27B NVFP4 model achieves 41% faster single-session inference and 23-25% faster concurrent inference on DGX Spark compared to Unsloth's version.

0 favorites 0 likes
#inference-speed

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model

Hugging Face Daily Papers · 2026-06-30 Cached

FlexiSLM introduces dynamic frame rate capabilities for speech input and output in spoken language models, outperforming fixed-frame-rate models and enabling controllable inference speed.

0 favorites 0 likes
#inference-speed

@sama: oh and also...750 token/sec coming to 5.6 sol in july!

X AI KOLs · 2026-06-26

Sam Altman announces that a model offering 750 tokens per second will be available for 5.6 SOL in July.

0 favorites 0 likes
#inference-speed

@agupta: I love that in this arms race for speed there is a seed stage YC company @wafer_ai fighting for the top spot

X AI KOLs Following · 2026-06-25 Cached

A seed-stage YC startup, Wafer AI, is competing in the inference speed race, with Databricks achieving 392 token/s on GLM-5.2, topping Artificial Analysis.

0 favorites 0 likes
#inference-speed

@Thom_Wolf: Multi-agents collaborations are among the most interesting agent behaviors right now! We did an experiment the other da…

X AI KOLs Timeline · 2026-06-25 Cached

An experiment with over 100 AI agents collaborating for a week to improve Gemma 4 inference speed in vLLM achieved a 5x speedup, revealing emergent behaviors like self-policing, division of labor, and communal knowledge sharing.

0 favorites 0 likes
#inference-speed

@HuggingPapers: Geometric Action Model for Robot Policy Learning Repurposes a geometric foundation model as one backbone for perception…

X AI KOLs Following · 2026-06-16 Cached

Geometric Action Model repurposes a geometric foundation model for robot policy learning, achieving 85.5% on LIBERO-Plus with 6.9 ms inference, 55× faster than baselines.

0 favorites 0 likes
#inference-speed

@rohanpaul_ai: Quite incredible, MiniMax Sparse Attention cuts attention compute by 28.4X at 1M tokens, with 14.2X faster prefill and …

X AI KOLs Following · 2026-06-15 Cached

MiniMax Sparse Attention (MSA) achieves up to 28.4x reduction in attention compute at 1M tokens by adding a routing branch that selectively chooses key-value blocks for attention, enabling 14.2x faster prefill and 7.6x faster decoding on H800 GPUs while matching full attention benchmark performance.

0 favorites 0 likes
#inference-speed

@charles_irl: Many are belatedly realizing that intelligence must be open. For open intelligence to succeed, developers must work tog…

X AI KOLs Following · 2026-06-15 Cached

A collaboration between Modal, SGLang, and Z Lab integrates DFlash speculation into SGLang, achieving up to 4.3x throughput improvement for Alibaba's Qwen 397B-A17B model, advancing open intelligence.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback