Tag
A blog post detailing performance optimizations in llama.cpp that make prompt lookup drafting up to 42x faster and reduce memory usage by 2.6x, based on techniques from Daniel Lemire and Martin Ankerl.
The article provides a step-by-step guide to tuning a server for benchmarking to reduce run-to-run noise and ensure measurements are repeatable, with techniques like hardware inspection and core pinning.
This article details the porting of the Kev-0.8B model, a Qwen3.5 variant, to Core ML, demonstrating substantial improvements in inference speed and memory efficiency on Apple Silicon devices through benchmarks.
This article describes how Anthropic tripled the speed of Claude.ai in two weeks, along with the author's experience applying this method to optimize their own website, and the release of the '闪电.skill' tool for public use.
This article discusses techniques for writing efficient C++ code, emphasizing data-oriented design and performance optimization for applications like games and real-time processing.
PSA advising to increase the -cram parameter in llama.cpp for better performance in agentic workflows with long contexts, based on personal experience with Qwen 27B 3.8.
The article presents an open-source study on optimizing latency for local vision-language models through benchmarking and techniques like native batching and MLX quantization, achieving significant speedups while maintaining decision accuracy on Apple hardware.
In a two-week sprint, the team used an internal Claude model to achieve a 3x speed improvement for key user journeys on claude.ai and the desktop app, significantly reducing wait times without any customer-facing incidents.
Adaptive Lossless floating-Point (ALP) Encoding is a new lightweight encoding for floating-point data in Apache Parquet, offering compression ratios similar to zstd with faster decompression, random-access support, and GPU/SIMD-friendly decoding.
Modal details how they optimized inference performance for trillion-parameter coding agents, achieving significant improvements in throughput and interactivity for their service.
The paper introduces a method to prune encoder layers in the Whisper ASR model, reducing encoder size by 18.5% and recovering performance through unlabeled data distillation, with code and a pre-trained model released for adoption.
Halogen version 0.12.0 fixes performance degradation at high context depths, showing improved decode and prefill speeds for Qwen3.8-Flash-Next at 1 million tokens of context on AMD Ryzen AI Max+ hardware.
This paper proposes Cache-to-Cache (C2C), a novel paradigm for direct semantic communication between large language models using KV-Cache, which enhances response quality and reduces latency compared to text-based methods.
This article explains how GPU memory is utilized during large language model (LLM) inference, breaking it down into four key components: model weights, KV cache, activations/workspace, and runtime overhead. It highlights the importance of quantization in optimizing memory usage for better performance.
Go 1.27 introduces size-specialized memory allocation for allocations of 80 bytes or fewer, improving allocation speeds by 20-30% and overall program performance by up to 1% for allocation-heavy code.
A technique to offload the KV cache of Qwen3.8-Flash-Next to system RAM is demonstrated, allowing long-context inference with minimal decode slowdown by leveraging the model's efficient architecture.
Polymarket has built in-house indexing systems powered by rindexer in Rust, achieving near-instant updates for trades, positions, and balances with onchain data events processed up to 28 seconds faster.
The co-inventor of ChatGPT announces the release of a new AI model named Jev, trained with RLCD, claiming it is 20-200x faster, 40-400x cheaper, and optimized for composable intelligence as a path to AGI.
An implementation of expert lookahead achieves over 10% performance improvement for MoE models running on low-memory devices using slotstream, with additional gains from a correction model.
This pull request adds missing AMD GCN MMQ configuration to ggml-cuda for HIP, enhancing prefill performance for RDNA2 GPUs such as MI50 and MI60 in the llama.cpp inference library.