performance-benchmarking

Tag

Cards List
#performance-benchmarking

GPT-6 and Opus 5.5's biggest revolution isn't performance, its speed and cost.

Reddit r/singularity · 19h ago

GPT-6 Sol and Claude Opus 5.5 achieve near-frontier performance at a fraction of the cost and speed of previous generations, highlighting major efficiency gains.

0 favorites 0 likes
#performance-benchmarking

Subnormal floating-point numbers are expensive… on Intel processors

Lobsters Hottest · 2026-09-15 Cached

This article benchmarks the performance of subnormal floating-point numbers on Intel, AMD, and ARM processors, revealing that Intel processors experience significant slowdowns with subnormals while AMD and ARM are unaffected.

0 favorites 0 likes
#performance-benchmarking

exllamav3 comfortably beats llama.cpp running CPU-offloaded Qwen-3.8-Flash-Next on my setup!

Reddit r/LocalLLaMA · 2026-09-07

A user benchmarks exllamav3 against llama.cpp for CPU-offloaded inference, showing exllamav3 is faster for Qwen models but slower for others, depending on hardware and model architecture.

0 favorites 0 likes
#performance-benchmarking

Question: Why is prefill unbelievably faster in vLLM than other inference engines?

Reddit r/LocalLLaMA · 2026-09-01

The user shares benchmark results showing vLLM's significantly faster prefill performance compared to llama.cpp and other engines, and questions the technical reasons behind this speed difference.

0 favorites 0 likes
#performance-benchmarking

FlashMLA sm_120 kernel build with 2-3x performance increase from SPDA

Reddit r/LocalLLaMA · 2026-08-30

The author built FlashMLA for consumer-grade Blackwell sm_120, achieving 2-3x performance gains over PyTorch SDPA in attention-heavy workloads like long-context training and sparse prefill.

0 favorites 0 likes
#performance-benchmarking

3 experiments running dsv4-flash-0731 q4+ quants on 128GB RAM + ~60 GB VRAM (with a quite bad pcie infra) with an acceptable tgs and relatively acceptable pp speed

Reddit r/LocalLLaMA · 2026-08-22

The author conducted experiments to run DeepSeek-V4-Flash-0731 with 4-bit quantizations on a 128GB RAM system, using optimizations like memory mlocking and prompt processing strategies to achieve acceptable inference speeds.

0 favorites 0 likes
#performance-benchmarking

New Google Gemma 4 12B Claims Near-26B Performance - We Tested Both!

Reddit r/LocalLLaMA · 2026-06-03

Google's new Gemma 4 12B model claims near-26B performance. In a local test on RTX 4090, the 26B-A4B model was faster and better but the 12B used less VRAM, making it suitable for laptops.

0 favorites 0 likes
← Back to home

Submit Feedback