performance-metrics

Tag

Cards List
#performance-metrics

HarnessTax (2 minute read)

TLDR AI · 3d ago Cached

The article introduces 'HarnessTax' and explores how much the harness, or development environment, affects the performance and effectiveness of coding agents.

0 favorites 0 likes
#performance-metrics

Speed is the wrong question, safety is one of the components of performance

Reddit r/artificial · 5d ago

The article argues that safety and alignment are integral to AI performance, and that slowing down isn't the issue—what matters is defining progress to include factors like obedience and honesty.

0 favorites 0 likes
#performance-metrics

15ms at P50 memory retrieval does absolutely nothing for a voice agent

Reddit r/AI_Agents · 6d ago

The article discusses key challenges in memory retrieval for voice agents, emphasizing the need for measuring P99 latency per turn, using prefetching, and budgeting memory tokens to reduce latency and improve user experience.

0 favorites 0 likes
#performance-metrics

Human Bench 3.0, Humanity-2 AI-0

Reddit r/singularity · 2026-09-09

Introduces Human Bench 3.0, a benchmark for evaluating AI systems in comparison to human performance, with Humanity-2 AI-0 likely indicating a specific score or version.

0 favorites 0 likes
#performance-metrics

Another qwen 3.8 27b showcase - gta style prompt - also a remainder to use ngram in your configs.

Reddit r/LocalLLaMA · 2026-08-20

A showcase of using the Qwen 3.8 27B model with 128k context to create a fully playable GTA Vice City-style game, including performance benchmarks and configuration tips.

0 favorites 0 likes
#performance-metrics

GLM 5.3 SlopCodeBench Results

Reddit r/LocalLLaMA · 2026-08-20

GLM 5.3 benchmark results on SlopCodeBench show it scoring 47.1% on one subset and tying with Fable 5 and GPT-5.6 Sol on another, demonstrating that all models struggle with the unsaturated coding benchmark.

0 favorites 0 likes
#performance-metrics

@elonmusk: Intelligence/Joule will keep improving

X AI KOLs Timeline · 2026-08-16 Cached

Elon Musk highlights continuous improvement in Intelligence/Joule, with Amjad Masad reporting an 18x increase in intelligence per joule over 16 months.

0 favorites 0 likes
#performance-metrics

Qwen3.8-27B abliterated FP8: refusal 64–99% → 0–6%, and MMLU/GSM8K move less than 1.3 points

Reddit r/LocalLLaMA · 2026-08-16

The article discusses the evaluation of the abliterated Qwen3.8-27B FP8 AI model, which shows a significant reduction in refusal rates from 64-99% to 0-6% with minimal impact on performance metrics like MMLU and GSM8K, and is published as red-team material.

0 favorites 0 likes
#performance-metrics

The mean means nothing

Lobsters Hottest · 2026-07-28 Cached

A blog post illustrating how relying solely on the mean can be misleading when evaluating performance improvements, using synthetic latency data to show the importance of looking at the full distribution via percentiles, density plots, and CDFs.

0 favorites 0 likes
#performance-metrics

A complementary study on PlanGPT: Evaluation with defined Performance Metrics and comparison with a planner

arXiv cs.AI · 2026-06-10 Cached

This paper presents a complementary evaluation of PlanGPT, a large language model for automated planning, using plan cost and plan generation time metrics, and finds that PlanGPT performs no better than a greedy search strategy.

0 favorites 0 likes
#performance-metrics

eTPS Site Plan – Simple Leaderboard + What You’ll Actually See

Reddit r/artificial · 2026-05-07

The author introduces the site plan for effectiveTPS, a tool designed to compare local AI models using a new 'effective TPS' metric alongside raw speed and latency. It aims to provide a simple leaderboard that highlights useful output quality over raw marketing numbers.

0 favorites 0 likes
← Back to home

Submit Feedback