Tag
This pull request adds missing AMD GCN MMQ configuration to ggml-cuda for HIP, enhancing prefill performance for RDNA2 GPUs such as MI50 and MI60 in the llama.cpp inference library.
The DSPy 3.4 release candidate is announced, moving away from litellm to improve import speed and reduce dependencies. Community members like Drew Breunig have praised the improvements.
A performance throttle in the Apple M3 Neural Engine was discovered when weight sizes are multiples of 1 MiB, causing throughput drops. By avoiding the problematic DMA path, token throughput for models like Llama 3.2 and Qwen3-8B was significantly improved.
An author improved MLX performance on an M3 Ultra Studio to achieve 73 tok/s for a 4-bit Qwen3.8-Flash-Next model, which shows intelligence scores matching GPT-5.6 Sol, highlighting the growing potential of local AI.
Swift-Qwen3.8-27B is UkisAI's reasoning-efficient derivative of Qwen3.8-27B, reducing thinking tokens by 58.3% while maintaining near-identical performance.
The authors have created a lightweight, performant asynchronous reinforcement learning stack in pure JAX, sharing a work log with insights on inference, RDMA weight transfer, memory optimizations, and sharding for scaling RL systems.
This article provides a comprehensive guide on Jump Labels in the Linux kernel, explaining their hardware background, usage, and implementation details for kernel programmers.
Michael Stapelberg details how he removed Debian Code Search's last cgo dependency by reimplementing the TurboPFor integer codec in Go using newly available SIMD support, leveraging AVX512 to exceed the reference implementation's performance.
The article details the merging of Multi-Token Prediction (MTP) support into ik_llama.cpp, which significantly enhances token generation speed for the Qwen3.8-Flash-Next model, with benchmarks showing up to double the speed on high-end hardware.
Research indicates that distilling AI models into architectures similar to the teacher model results in better performance compared to distilling into very different architectures.
A Rustdoc team member achieved a 33% performance improvement in Rustdoc through a series of optimizations and bug fixes, addressing a recursion limit issue that affected documentation generation.
RespanAI introduces Respan, a solution designed to handle the performance degradation loop in AI products, accompanied by a promotional film.
The article explains that CPUs optimistically implement memory ordering and rely on rollback mechanisms for contention, clarifying misconceptions about strongly versus weakly ordered architectures.
mold is a massively parallel linker that reduces link times by applying data parallelism across all passes, achieving 2.4–16.1x faster performance than state-of-the-art linkers like lld.
Discusses solutions to the 1+N query problem, a common performance issue in database interactions, often related to ORM usage.
A user fixed the MTP head on the Ornith1.5 35B A3B model, achieving a 33% reduction in wall clock time for tasks, making it significantly faster for local HAM radio applications.
The Ornith 1.5 35B A3B model's MTP tensors appear to be uninitialized, causing poor speculative decoding performance, and grafting the trained head from Qwen3.6-35B-A3B improves speed by 29%.
Datadog rebuilt their Git serving infrastructure with a custom mirror called gitretriever, handling 20 times more CI traffic while maintaining low latency and reducing CPU usage.
This paper integrates SmoothQuant into PyTorch's native stack for efficient INT8 inference of small NLP models on Intel Xeon CPUs, achieving up to 5.8× speedup with negligible accuracy loss.
This article presents performance results and optimization tips for using llama.cpp with RPC to run AI models across a 5070 Ti and 1080 Ti GPUs over ethernet, emphasizing trade-offs between prefill and token generation speeds.