performance-optimization

Tag

Cards List
#performance-optimization

Linux 7.3 improves performance when running out of vRAM

Hacker News Top ↗ · 2026-08-18 Cached

Linux 7.3 merges kernel patches that improve VRAM management, enhancing performance stability in games when physical VRAM is exceeded through optimized memory eviction and caching.

0 favorites 0 likes
#performance-optimization

Llama.cpp v0.1.0

Hacker News Top ↗ · 2026-08-17 Cached

Llama.cpp v0.1.0 is a C/C++ implementation for efficient LLM and VLM inference, supporting a wide range of hardware with minimal setup and high performance.

0 favorites 0 likes
#performance-optimization

Qwen3.8-27B is now up to ~3× faster on Apple Silicon with mlx-dspark

Reddit r/LocalLLaMA ↗ · 2026-08-14

mlx-dspark v0.10.0 adds support for Qwen3.8-27B on Apple Silicon, providing up to 3x faster inference through speculative decoding with lossless verification.

0 favorites 0 likes
#performance-optimization

NInfer day0 support for Qwen3.8 27b: ~200 tok/s generation, with tons of engine improvments

Reddit r/LocalLLaMA ↗ · 2026-08-14

NInfer adds day-0 support for the Qwen3.8-27B model, achieving around 200 tokens per second on a single RTX 5090 with speculative decoding, and includes engine improvements like concurrent requests and kernel optimizations.

0 favorites 0 likes
#performance-optimization

@PyTorch: AMD has been upstreaming optimizations for improved FP8 training support in PyTorch/TorchTitan and PyTorch/TorchAO, mak…

X AI KOLs Timeline ↗ · 2026-08-13 Cached

AMD upstreamed optimizations to PyTorch/TorchTitan and TorchAO for FP8 training on AMD Instinct GPUs, achieving up to 13.4% throughput gains on Llama3-8B and recovering 89% of FP8 quantization overhead on DeepSeek-V3 via fused Triton kernels.

0 favorites 0 likes
#performance-optimization

@omarsar0: TokTier makes tokenization stateful for agentic serving. Across 153,951 real agent calls with a 94.1% prompt-cache hit …

X AI KOLs Following ↗ · 2026-08-03 Cached

TokTier is a stateful tokenization service that cuts tokenization overhead in agentic LLM serving by reusing stable token boundaries and running exact CPU/GPU tokenization, reducing time-to-first-token by 16–34% under vLLM.

0 favorites 0 likes
#performance-optimization

I built an open-source tool that reduces TensorBoard trace sizes by 90%+ for JAX/XLA (XProf Cubism Reducer)

Reddit r/LocalLLaMA ↗ · 2026-07-30

An open-source tool called XProf Cubism Reducer reduces TensorBoard trace sizes by over 90% for JAX/XLA, making performance profiling more efficient.

0 favorites 0 likes
#performance-optimization

@akshay_pachaar: every inference engine makes the same mistake. an inference engine like vLLM or SGLang is the software sitting between …

X AI KOLs Following ↗ · 2026-07-27 Cached

LMCache is an open-source KV cache management layer that separates cache I/O from compute, plugging into vLLM, SGLang, and TensorRT-LLM to achieve up to 14x faster time-to-first-token and 4x faster decoding by parallelizing cache lookups and sharing GPU memory.

0 favorites 0 likes
#performance-optimization

@GergelyOrosz: That @Sirupsen measured, published then memorized about 3x these numbers on networking+latency "basics" is only surpris…

X AI KOLs Following ↗ · 2026-07-21 Cached

Gergely Orosz praises Sirupsen for his deep understanding of networking and latency basics, referencing an interview where Sirupsen used napkin math to build a 10x cheaper or faster product.

0 favorites 0 likes
#performance-optimization

@Alacritic_Super: If you are building production LLM applications, learn LLM Caching. Caching can reduce latency, GPU utilization, and AP…

X AI KOLs Timeline ↗ · 2026-07-12 Cached

This article emphasizes the importance of LLM caching in production systems to reduce latency, GPU utilization, and costs, and introduces LMCache, an open-source KV cache management layer for scalable LLM inference.

0 favorites 0 likes
#performance-optimization

@PyTorch: https://bit.ly/4yawNqB..*

X AI KOLs Timeline ↗ · 2026-07-10 Cached

This blog post from PyTorch presents novel kernel fusion techniques for normalization ops like LayerNorm and RMSNorm, achieving significant speedups by reducing memory-IO overhead. Techniques include Lazy Pre-Norm and Multi-CTA Norm Fusion, hiding up to 90% of normalization latency when fused with GEMMs, and the FlashNormAttention algorithm achieving up to 35% kernel speedup.

0 favorites 0 likes
#performance-optimization

@ariG23498: I have always admired @stevhliu's work. I consider his technical writeups to be among the best there is. In the latest …

X AI KOLs Timeline ↗ · 2026-07-08 Cached

A thread highlighting a technical blog series on how Hugging Face's transformers library loads models efficiently, covering meta device, safetensors, CUDA caching, and more.

0 favorites 0 likes
#performance-optimization

GLM-5.2 on 8xB200: the deployment math nobody spells out - NVFP4 + 2x TP=4 replicas should beat TP=8 by ~2x. Full config guidance inside.

Reddit r/LocalLLaMA ↗ · 2026-07-07

The article provides the optimal deployment configuration for GLM-5.2 on 8xB200 nodes, showing that NVFP4 with two TP=4 replicas achieves roughly 2x throughput over FP8 TP=8, with detailed performance data and caveats.

0 favorites 0 likes
#performance-optimization

@superalesha: I sped up deepseek v4 flash by 29x on my 4x3090s !!! No, its not joke. 15 -> 443 t/s. a 23k prompt used to take 25 mins…

X AI KOLs Timeline ↗ · 2026-07-07 Cached

A user achieved a 29x speedup for DeepSeek V4 Flash inference on 4x RTX 3090 GPUs by optimizing llama.cpp, reducing a 23k prompt from 25 minutes to 53 seconds.

0 favorites 0 likes
#performance-optimization

Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?

Hugging Face Daily Papers ↗ · 2026-07-01 Cached

This paper audits three performance-optimization benchmarks (GSO, SWE-Perf, SWE-efficiency) for coding agents, finding that runtime instability, scoring rules, and task coverage significantly affect reliability, and that many tasks are already solved by at least one public submission.

0 favorites 0 likes
#performance-optimization

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving

Hugging Face Daily Papers ↗ · 2026-07-01 Cached

ELDR is an expert-locality-aware decode router for prefill-decode disaggregated Mixture-of-Experts serving that predicts expert activations from prefill signatures and routes requests to minimize latency, implemented in vLLM and achieving 5.9-13.9% reduction in median TPOT on up to 40 GPUs.

0 favorites 0 likes
#performance-optimization

Superpowers 6

Hacker News Top ↗ · 2026-06-30 Cached

Superpowers 6大幅提升了开发速度和成本效率,通过Fable的优化实现最高50%更快构建和60%更低token消耗,同时改进了对多个AI模型和编码代理的支持。

0 favorites 0 likes
#performance-optimization

Faster KNN search in Manticore: 2-pass HNSW, batched distances, and AVX-512

Hacker News Top ↗ · 2026-06-26 Cached

Manticore's KNN search gets up to 29% faster with 2-pass HNSW, batched distances, compile-time distance specialization, and AVX-512 support.

0 favorites 0 likes
#performance-optimization

EGG: An Expert-Guided Agent Framework for Kernel Generation

arXiv cs.AI ↗ · 2026-06-26 Cached

EGG is an expert-guided agent framework that decomposes GPU kernel generation into algorithmic structure design and hardware-specific tuning, using a stage-aware multi-agent collaboration mechanism. It achieves a 2.13x average speedup over PyTorch on KernelBench and real-world workloads.

0 favorites 0 likes
#performance-optimization

@poteto: https://x.com/poteto/status/2069824386283319343

X AI KOLs Following ↗ · 2026-06-24 Cached

The article draws parallels between managing engineering teams and managing AI agents, using Andy Grove's management principles to build reliable agent loops, illustrated through a performance debugging case study at Cursor.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback