throughput

Tag

Cards List
#throughput

Higher acceptance length, slower prose: Ling’s n=1/2/3 MTP test on one Spark

Reddit r/LocalLLaMA · 5d ago

Benchmark tests on Ling-3.0-flash show that higher acceptance length in multi-token prediction decreases prose throughput, with n=1 being the most efficient setting for the evaluated workloads.

0 favorites 0 likes
#throughput

Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning

Hugging Face Daily Papers · 2026-09-03 Cached

The paper proposes Random Attention, a KV cache eviction method that uses random selection instead of scoring, matching selective methods while improving throughput in reasoning tasks.

0 favorites 0 likes
#throughput

NVIDIA reports up to 30× more agentic throughput per MW on Vera Rubin—but tokens/MW still is not completed work/MW

Reddit r/ArtificialInteligence · 2026-08-26

NVIDIA reports up to 30× higher agentic throughput per megawatt with Vera Rubin compared to GB300, but emphasizes that production benchmarks should include additional metrics beyond tokens per watt.

0 favorites 0 likes
#throughput

Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

NVIDIA Blog · 2026-08-24 Cached

NVIDIA's Vera Rubin NVL72 system demonstrates up to 30x higher throughput per megawatt for agentic AI workloads, setting a new efficiency standard for AI infrastructure.

0 favorites 0 likes
#throughput

Wi-Fi 8 is the first wireless upgrade in years that isn't chasing speed

Hacker News Top · 2026-08-23 Cached

Wi-Fi 8 shifts focus from speed to reliability, aiming to reduce interference and latency while maintaining similar data rates as Wi-Fi 7.

0 favorites 0 likes
#throughput

@maximelabonne: Nature healed: LFM2.5-2.6B is free on OpenRouter and has good throughput! Try it in your favorite agentic harness today.

X AI KOLs Following · 2026-08-21 Cached

The LFM2.5-2.6B AI model is free to use on OpenRouter with good throughput, and is promoted for agentic harness applications.

0 favorites 0 likes
#throughput

I measured the 3 claims Users in this Sub all handed me on the last local-agent post. One of you out-predicted my own hypothesis. Learn It All not Know It All rules

Reddit r/AI_Agents · 2026-08-19

The user tested scaling local AI agents with a Qwen 27B model, finding that adding more agents increases throughput only up to a point due to memory bandwidth limits, with long prompts benefiting more from parallelism.

0 favorites 0 likes
#throughput

DFlash 2: Keep Drafting Parallel

Reddit r/LocalLLaMA · 2026-08-18 Cached

DFlash 2 improves speculative decoding by predicting tokens in parallel, achieving over 20% more output per verification pass with minimal latency, and is integrated into major inference engines like SGLang and vLLM.

0 favorites 0 likes
#throughput

A 124B emitted 15,128 tokens in a single response on one DGX Spark, decode went 35.62 → 35.68 tok/s across the whole thing

Reddit r/LocalLLaMA · 2026-08-13

An observation of Ling-3.0-flash (124B) on one DGX Spark generating 15,128 tokens in a single response with stable decode throughput around 35.6 tok/s, highlighting long-context decoding performance.

0 favorites 0 likes
#throughput

Concurrency vs. Throughput: why more parallelism can make databases slower

Lobsters Hottest · 2026-08-07 Cached

A PlanetScale engineering post analyzes a MySQL outage caused by a long-running transaction and high concurrency, explaining how parallelism can degrade throughput and how Vitess's transaction pool handled (and amplified) the issue.

0 favorites 0 likes
#throughput

@TeachTheMachine: Measuring Performance of Transformer Inference

X AI KOLs Timeline · 2026-08-06 Cached

A practical tutorial on measuring transformer inference performance, covering metrics like latency, TTFT, throughput, memory usage, and benchmarking techniques for LLMs.

0 favorites 0 likes
#throughput

Intern S2 Mobius

Reddit r/LocalLLaMA · 2026-08-05

Intern-S2 Mobius is a Qwen3.5-35B derived model with an architectural difference claimed to improve throughput and reduce token consumption.

0 favorites 0 likes
#throughput

@wafer_ai: BREAKING We're on Hacker News again we figured out how to serve Kimi K3 at 3.8x higher throughput and 71% lower cost on…

X AI KOLs Following · 2026-08-02 Cached

Wafer announces that it can serve Kimi K3 on AMD MI355X at 3.8x higher throughput and 71% lower cost than on B200 nodes, arguing that AMD's large VRAM and software support make it the best performance-per-dollar choice for frontier models.

0 favorites 0 likes
#throughput

@FinanceYF5: Kimi K3 went viral in two days. Deedy says there's a whole set of compute economics behind it. 1/ K3 has only been online for two days, but it has already reached #10 on OpenRouter, processing 140 billion tokens daily. However, the servers can't handle the load: throughput dropped from 30 Token/s to 13, latency spiked to 72 seconds, and time to first token exceeded 20 seconds...

X AI KOLs Following · 2026-07-20 Cached

Kimi K3 reached #10 on OpenRouter within two days of launch, processing 140 billion tokens daily, but high load caused throughput to drop and latency to spike.

0 favorites 0 likes
#throughput

@NVIDIAAI: As AI models continue to grow in scale and capability, shaping a model matters just as much as its size. We're introduc…

X AI KOLs Timeline · 2026-07-13 Cached

NVIDIA introduces a series on AI Model Co-Design, explaining how model dimensions affect GPU performance and the trade-offs between throughput and interactivity for LLM deployment. The first post provides a practical primer on designing hardware-friendly LLMs to improve system throughput and user responsiveness.

0 favorites 0 likes
#throughput

The 4-Bitter Lesson: Balancing Stability and Performance in NVFP4 RL

Hacker News Top · 2026-07-10 Cached

This article presents a recipe for low-precision (NVFP4) RL training that balances throughput and stability, addressing issues from forward and backward pass quantization errors.

0 favorites 0 likes
#throughput

@WescheNex1q: 16 people chatting with Qwen3.6-35B at once ONE DGX Spark. This is a real capture, not a mockup: every token you see re…

X AI KOLs Timeline · 2026-07-09 Cached

A real-time demo shows 16 concurrent users chatting with Qwen3.6-35B on a single DGX Spark, achieving peak 440 tok/s total and 105 tok/s per user using NVFP4 + MTP-3 on vLLM.

0 favorites 0 likes
#throughput

@YRSM_Simon: Crazy

X AI KOLs Timeline · 2026-07-09 Cached

DeepSeek-V4-Flash-DSpark achieves 328 tok/s single inference and 1.7k tok/s batch throughput on 4x RTX PRO 6000 GPUs.

0 favorites 0 likes
#throughput

@Alacritic_Super: If you want to master LLM inference, start with these three papers. They introduced many of the ideas powering today's …

X AI KOLs Timeline · 2026-07-08 Cached

This thread recommends three key papers for mastering LLM inference: PagedAttention, Sarathi-Serve, and SGLang, which introduce efficient memory management, chunked prefills, and structured generation techniques used in modern inference engines like vLLM and TensorRT-LLM.

0 favorites 0 likes
#throughput

Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding

Hugging Face Daily Papers · 2026-07-07 Cached

The paper introduces Nemotron-Labs-Diffusion, a tri-mode language model that unifies autoregressive, diffusion, and self-speculation decoding, achieving superior throughput and efficiency compared to existing models.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback