performance-optimization

Tag

Cards List
#performance-optimization

ggml-cuda: hip: add missing AMD GCN MMQ config by thelittlefireman · Pull Request #27841 · ggml-org/llama.cpp - PP improvements for RDNA2(MI50, MI60)

Reddit r/LocalLLaMA ↗ · 2026-09-12 Cached

This pull request adds missing AMD GCN MMQ configuration to ggml-cuda for HIP, enhancing prefill performance for RDNA2 GPUs such as MI50 and MI60 in the llama.cpp inference library.

0 favorites 0 likes
#performance-optimization

@MaximeRivest: DSPy 3.4 release candidate it out, we are moving off litellm. It always bothered me how slow dspy was to import and how…

X AI KOLs Following ↗ · 2026-09-11 Cached

The DSPy 3.4 release candidate is announced, moving away from litellm to improve import speed and reduce dependencies. Community members like Drew Breunig have praised the improvements.

0 favorites 0 likes
#performance-optimization

Getting 50 GB/S Back from the Apple Neural Engine

Hacker News Top ↗ · 2026-09-10 Cached

A performance throttle in the Apple M3 Neural Engine was discovered when weight sizes are multiples of 1 MiB, causing throughput drops. By avoiding the problematic DMA path, token throughput for models like Llama 3.2 and Qwen3-8B was significantly improved.

0 favorites 0 likes
#performance-optimization

@ashxhart: Been giving my M3 Ultra Studio some love and tinkering with MLX. Started the afternoon at 42 tok/s. Now: 73 tok/s, and …

X AI KOLs Timeline ↗ · 2026-09-08 Cached

An author improved MLX performance on an M3 Ultra Studio to achieve 73 tok/s for a 4-bit Qwen3.8-Flash-Next model, which shows intelligence scores matching GPT-5.6 Sol, highlighting the growing potential of local AI.

0 favorites 0 likes
#performance-optimization

ukisai/Swift-Qwen3.8-27b

Hugging Face Models Trending ↗ · 2026-09-08 Cached

Swift-Qwen3.8-27B is UkisAI's reasoning-efficient derivative of Qwen3.8-27B, reducing thinking tokens by 58.3% while maintaining near-identical performance.

0 favorites 0 likes
#performance-optimization

@_DivyaMakkar: Spent some time creating a lightweight, performant async RL stack in pure JAX with @AdityaMakkar0! We share a work log …

X AI KOLs Following ↗ · 2026-09-07

The authors have created a lightweight, performant asynchronous reinforcement learning stack in pure JAX, sharing a work log with insights on inference, RDMA weight transfer, memory optimizations, and sharding for scaling RL systems.

0 favorites 0 likes
#performance-optimization

What every kernel programmer should know about Jump Labels

Lobsters Hottest ↗ · 2026-09-07 Cached

This article provides a comprehensive guide on Jump Labels in the Linux kernel, explaining their hardware background, usage, and implementation details for kernel programmers.

0 favorites 0 likes
#performance-optimization

Debian Code Search: Fast TurboPFor with Go SIMD

Lobsters Hottest ↗ · 2026-09-06 Cached

Michael Stapelberg details how he removed Debian Code Search's last cgo dependency by reimplementing the TurboPFor integer codec in Go using newly available SIMD support, leveraging AVX512 to exceed the reference implementation's performance.

0 favorites 0 likes
#performance-optimization

Qwen3.8-Flash-Next MTP merged in ik_llama.cpp (integrated head or separate -md file)... 45 → 90 tok/s on a 5090 + 128GB, works down to a 12GB 4070

Reddit r/LocalLLaMA ↗ · 2026-09-03

The article details the merging of Multi-Token Prediction (MTP) support into ik_llama.cpp, which significantly enhances token generation speed for the Qwen3.8-Flash-Next model, with benchmarks showing up to double the speed on high-end hardware.

0 favorites 0 likes
#performance-optimization

@seclink: We have found that distilling the model to a target model with a highly similar architecture yields performance far sup…

X AI KOLs Following ↗ · 2026-09-01 Cached

Research indicates that distilling AI models into architectures similar to the teacher model results in better performance compared to distilling into very different architectures.

0 favorites 0 likes
#performance-optimization

How I made Rustdoc 33% faster in one week

Lobsters Hottest ↗ · 2026-08-28 Cached

A Rustdoc team member achieved a 33% performance improvement in Rustdoc through a series of optimizations and bug fixes, addressing a recursion limit issue that affected documentation generation.

0 favorites 0 likes
#performance-optimization

@RespanAI: Most AI products don’t fail. They just lose a little performance everywhere. We built Respan for that entire loop, and …

X AI KOLs Following ↗ · 2026-08-27 Cached

RespanAI introduces Respan, a solution designed to handle the performance degradation loop in AI products, accompanied by a promotional film.

0 favorites 0 likes
#performance-optimization

Memory ordering in CPUs

Lobsters Hottest ↗ · 2026-08-26 Cached

The article explains that CPUs optimistically implement memory ordering and rely on rollback mechanisms for contention, clarifying misconceptions about strongly versus weakly ordered architectures.

0 favorites 0 likes
#performance-optimization

mold: A Massively Parallel Linker

Lobsters Hottest ↗ · 2026-08-26 Cached

mold is a massively parallel linker that reduces link times by applying data parallelism across all passes, achieving 2.4–16.1x faster performance than state-of-the-art linkers like lld.

0 favorites 0 likes
#performance-optimization

Solving the 1+N query problem

Lobsters Hottest ↗ · 2026-08-25

Discusses solutions to the 1+N query problem, a common performance issue in database interactions, often related to ORM usage.

0 favorites 0 likes
#performance-optimization

Fixed the MTP head on Ornith1.5 35B A3B. +3% TPS -33% wall clock

Reddit r/LocalLLaMA ↗ · 2026-08-22

A user fixed the MTP head on the Ornith1.5 35B A3B model, achieving a 33% reduction in wall clock time for tasks, making it significantly faster for local HAM radio applications.

0 favorites 0 likes
#performance-optimization

If you are wondering why Ornith 1.5 35B A3B with MTP is so slow, this is why

Reddit r/LocalLLaMA ↗ · 2026-08-20 Cached

The Ornith 1.5 35B A3B model's MTP tensors appear to be uninitialized, causing poor speculative decoding performance, and grafting the trained head from Qwen3.6-35B-A3B improves speed by 29%.

0 favorites 0 likes
#performance-optimization

20× the CI traffic without getting slower: How we rebuilt Git serving at Datadog

Lobsters Hottest ↗ · 2026-08-20 Cached

Datadog rebuilt their Git serving infrastructure with a custom mirror called gitretriever, handling 20 times more CI traffic while maintaining low latency and reducing CPU usage.

0 favorites 0 likes
#performance-optimization

Efficient INT8 Inference of Small NLP Models on Server CPUs with PyTorch Native Stack

arXiv cs.CL ↗ · 2026-08-20 Cached

This paper integrates SmoothQuant into PyTorch's native stack for efficient INT8 inference of small NLP models on Intel Xeon CPUs, achieving up to 5.8× speedup with negligible accuracy loss.

0 favorites 0 likes
#performance-optimization

Don't ignore llama.cpp RPC with old hardware. Results of a 5070 Ti and 1080 Ti over gigabit ethernet: it's actually functional.

Reddit r/LocalLLaMA ↗ · 2026-08-18

This article presents performance results and optimization tips for using llama.cpp with RPC to run AI models across a 5070 Ti and 1080 Ti GPUs over ethernet, emphasizing trade-offs between prefill and token generation speeds.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback