performance-optimization

Tag

Cards List
#performance-optimization

42x Faster Prompt Lookup Drafting in llama.cpp

Reddit r/LocalLLaMA ↗ · 3d ago Cached

A blog post detailing performance optimizations in llama.cpp that make prompt lookup drafting up to 42x faster and reduce memory usage by 2.6x, based on techniques from Daniel Lemire and Martin Ankerl.

0 favorites 0 likes
#performance-optimization

Tuning a Server for Benchmarking

Lobsters Hottest ↗ · 3d ago Cached

The article provides a step-by-step guide to tuning a server for benchmarking to reduce run-to-run noise and ensure measurements are repeatable, with techniques like hardware inspection and core pinning.

0 favorites 0 likes
#performance-optimization

@Alex_tra_memory: Kev on Core ML After seeing this tweet, We got inspired to port Kev-0.8B (the Qwen3.5 variant) and ran Guess Who over 8…

X AI KOLs Timeline ↗ · 4d ago Cached

This article details the porting of the Kev-0.8B model, a Qwen3.5 variant, to Core ML, demonstrating substantial improvements in inference speed and memory efficiency on Apple Silicon devices through benchmarks.

0 favorites 0 likes
#performance-optimization

@AlchainHust: https://x.com/AlchainHust/status/2103711280364240936

X AI KOLs Timeline ↗ · 4d ago Cached

This article describes how Anthropic tripled the speed of Claude.ai in two weeks, along with the author's experience applying this method to optimize their own website, and the release of the '闪电.skill' tool for public use.

0 favorites 0 likes
#performance-optimization

Writing Efficient C++ Code

Hacker News Top ↗ · 4d ago Cached

This article discusses techniques for writing efficient C++ code, emphasizing data-oriented design and performance optimization for applications like games and real-time processing.

0 favorites 0 likes
#performance-optimization

PSA: llama.cpp -cram should be increased for agentic workflows (default is 8192)

Reddit r/LocalLLaMA ↗ · 5d ago

PSA advising to increase the -cram parameter in llama.cpp for better performance in agentic workflows with long contexts, based on personal experience with Qwen 27B 3.8.

0 favorites 0 likes
#performance-optimization

@yoheinakajima: glance-vlm speedlab is now open source! read: https://glance.yohei.me/speed/ try: https://github.com/yoheinakajima/glan…

X AI KOLs Timeline ↗ · 6d ago Cached

The article presents an open-source study on optimizing latency for local vision-language models through benchmarking and techniques like native batching and MLX quantization, achieving significant speedups while maintaining decision accuracy on Apple hardware.

0 favorites 0 likes
#performance-optimization

Once Claude can measure something, it can make it faster

Hacker News Top ↗ · 6d ago Cached

In a two-week sprint, the team used an internal Claude model to achieve a 3x speed improvement for key user journeys on claude.ai and the desktop app, significantly reducing wait times without any customer-facing incidents.

0 favorites 0 likes
#performance-optimization

Adaptive Lossless Floating-Point Encoding in Apache Parquet

Lobsters Hottest ↗ · 6d ago Cached

Adaptive Lossless floating-Point (ALP) Encoding is a new lightweight encoding for floating-point data in Apache Parquet, offering compression ratios similar to zstd with faster decompression, random-access support, and GPU/SIMD-friendly decoding.

0 favorites 0 likes
#performance-optimization

@charles_irl: After this, people kept asking us how @modal is able to serve agent inference so well. So we wrote it all down. New blo…

X AI KOLs Timeline ↗ · 6d ago Cached

Modal details how they optimized inference performance for trillion-parameter coding agents, achieving significant improvements in throughput and interactivity for their service.

0 favorites 0 likes
#performance-optimization

Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery

Hugging Face Daily Papers ↗ · 2026-09-23 Cached

The paper introduces a method to prune encoder layers in the Whisper ASR model, reducing encoder size by 18.5% and recovering performance through unlabeled data distillation, with code and a pre-trained model released for adoption.

0 favorites 0 likes
#performance-optimization

Qwen3.8-Flash-Next at 1M context on Strix Halo: 38 tok/s decode, 18 min prefill (halogen 0.12.0)

Reddit r/LocalLLaMA ↗ · 2026-09-19

Halogen version 0.12.0 fixes performance degradation at high context depths, showing improved decode and prefill speeds for Qwen3.8-Flash-Next at 1 million tokens of context on AMD Ryzen AI Max+ hardware.

0 favorites 0 likes
#performance-optimization

Cache-to-Cache: Direct Semantic Communication Between LLMs (2025)

Hacker News Top ↗ · 2026-09-18 Cached

This paper proposes Cache-to-Cache (C2C), a novel paradigm for direct semantic communication between large language models using KV-Cache, which enhances response quality and reduces latency compared to text-based methods.

0 favorites 0 likes
#performance-optimization

@akshay_pachaar: where does all the VRAM go during LLM inference? (4 ways GPU memory is used) loading the model is only the first part o…

X AI KOLs Timeline ↗ · 2026-09-17 Cached

This article explains how GPU memory is utilized during large language model (LLM) inference, breaking it down into four key components: model weights, KV cache, activations/workspace, and runtime overhead. It highlights the importance of quantization in optimizing memory usage for better performance.

0 favorites 0 likes
#performance-optimization

Size-Specialized Memory Allocation

Hacker News Top ↗ · 2026-09-17 Cached

Go 1.27 introduces size-specialized memory allocation for allocations of 80 bytes or fewer, improving allocation speeds by 20-30% and overall program performance by up to 1% for allocation-heavy code.

0 favorites 0 likes
#performance-optimization

You can offload most of Qwen3.8-Flash-Next's KV cache to RAM with little decode slowdown

Reddit r/LocalLLaMA ↗ · 2026-09-16

A technique to offload the KV cache of Qwen3.8-Flash-Next to system RAM is demonstrated, allowing long-context inference with minimal decode slowdown by leveraging the model's efficient architecture.

0 favorites 0 likes
#performance-optimization

@devjoshstevens: We have built our Polymarket indexing systems in-house, powered by rindexer in Rust. We now see onchain data events up …

X AI KOLs Timeline ↗ · 2026-09-15

Polymarket has built in-house indexing systems powered by rindexer in Rust, achieving near-instant updates for trades, positions, and balances with onchain data events processed up to 28 seconds faster.

0 favorites 0 likes
#performance-optimization

@CompleteSkeptic: After co-inventing ChatGPT, I kept asking myself: why have superhuman chat models not led to AGI? I’ve spent the last 2…

X AI KOLs Following ↗ · 2026-09-15 Cached

The co-inventor of ChatGPT announces the release of a new AI model named Jev, trained with RLCD, claiming it is 20-200x faster, 40-400x cheaper, and optimized for composable intelligence as a path to AGI.

0 favorites 0 likes
#performance-optimization

10%+ performance improvement on MoE ssd-streaming with expert-lookahead

Reddit r/LocalLLaMA ↗ · 2026-09-15

An implementation of expert lookahead achieves over 10% performance improvement for MoE models running on low-memory devices using slotstream, with additional gains from a correction model.

0 favorites 0 likes
#performance-optimization

ggml-cuda: hip: add missing AMD GCN MMQ config by thelittlefireman · Pull Request #27841 · ggml-org/llama.cpp - PP improvements for RDNA2(MI50, MI60)

Reddit r/LocalLLaMA ↗ · 2026-09-12 Cached

This pull request adds missing AMD GCN MMQ configuration to ggml-cuda for HIP, enhancing prefill performance for RDNA2 GPUs such as MI50 and MI60 in the llama.cpp inference library.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback