performance-testing

Tag

Cards List
#performance-testing

@nicebabycat: https://x.com/nicebabycat/status/2091726637155103126

X AI KOLs Following · yesterday Cached

This article provides a detailed test of the local deployment and performance of the Ling-3.0-tiny model on an Apple M5 chip Mac, demonstrating the feasibility of running a 7.9B parameter model at 47 tokens per second without a discrete GPU.

0 favorites 0 likes
#performance-testing

New/Old benchmark that provides a lot of answers for local LLM

Reddit r/LocalLLaMA · 3d ago

The article presents a benchmark tool for evaluating local LLM configurations, focusing on VRAM usage, performance metrics, and hardware optimization to assist developers in optimizing setups.

0 favorites 0 likes
#performance-testing

Qwen 3.8 27B KV f16 vs q8_0 are not equivalents

Reddit r/LocalLLaMA · 5d ago

A user shares test results indicating that KV cache types f16 and q8_0 are not equivalent for the Qwen 3.8 27B model, with f16 showing better detail and consistency, and provides configuration details for AMD ROCm hardware.

0 favorites 0 likes
#performance-testing

I measured the 3 claims Users in this Sub all handed me on the last local-agent post. One of you out-predicted my own hypothesis. Learn It All not Know It All rules

Reddit r/AI_Agents · 5d ago

The user tested scaling local AI agents with a Qwen 27B model, finding that adding more agents increases throughput only up to a point due to memory bandwidth limits, with long prompts benefiting more from parallelism.

0 favorites 0 likes
#performance-testing

I measured whether 2 local agents hitting 1 model run in parallel or just take turns. Batching is real, but it is not free using QWEN 3.8 27B 4bit on my MacBook Pro M3Max 128 GB Unified Memory 40 Core GPU

Reddit r/AI_Agents · 2026-08-18

The author experimented with two local agents running in parallel on a MacBook Pro M3Max using the QWEN 3.8 27B 4bit model, finding that batching enables concurrent execution but increases latency, with an optimal agent count around 4.

0 favorites 0 likes
#performance-testing

Curl Performance

Lobsters Hottest · 2026-08-14 Cached

Daniel Stenberg announces a new performance test suite for curl, with automated builds and results published publicly at curl.se/perf.

0 favorites 0 likes
#performance-testing

Geekbench 7 will push your computer or phone even harder for better benchmarking

The Verge · 2026-07-23 Cached

Primate Labs releases Geekbench 7 with new video/audio encoding tests, a redesigned multi-core test, and larger datasets for more accurate benchmarking of modern hardware.

0 favorites 0 likes
#performance-testing

6x MI50's on PCIE vs 4x MI50's on PEX8749 and 2x on PCIE

Reddit r/LocalLLaMA · 2026-07-11

A user benchmarks AMD MI50 GPUs across different PCIe configurations on an older X99 motherboard, comparing direct PCIe connections vs using a PEX8749 switch. Results show minimal performance difference with slight improvement in token generation speed.

0 favorites 0 likes
#performance-testing

Moby Dick Workout (2022)

Hacker News Top · 2026-07-05 Cached

Describes a performance test using the full text of Moby Dick to evaluate todo list and productivity apps, with test files provided for different app formats.

0 favorites 0 likes
#performance-testing

Evaluating Spec CPU2026

Hacker News Top · 2026-05-23 Cached

An in-depth evaluation of the new SPEC CPU2026 benchmark suite, which replaces SPEC CPU2017 with 52 workloads and a slower reference system (Ampere eMAG 8180), showing performance comparisons between modern CPUs.

0 favorites 0 likes
#performance-testing

Getting a feel for how fast X tokens/second really is.

Reddit r/LocalLLaMA · 2026-05-10

The author introduces a web-based script designed to help users intuitively understand token-per-second speeds in local LLM setups by simulating text, code, and reasoning generation rates.

0 favorites 0 likes
#performance-testing

Tested how OpenCode Works with SelfHosted LLMS: Qwen 3.5, 3.6, Gemma 4, Nemotron 3, GLM-4.7 Flash - v2

Reddit r/LocalLLaMA · 2026-04-22

A developer benchmarked multiple self-hosted LLMs (Qwen 3.5/3.6, Gemma 4, Nemotron 3, GLM-4.7) with OpenCode on two coding tasks, revealing speed and quality trade-offs on RTX 4080 hardware.

0 favorites 0 likes
← Back to home

Submit Feedback