Tag
The article introduces 'HarnessTax' and explores how much the harness, or development environment, affects the performance and effectiveness of coding agents.
The article argues that safety and alignment are integral to AI performance, and that slowing down isn't the issue—what matters is defining progress to include factors like obedience and honesty.
The article discusses key challenges in memory retrieval for voice agents, emphasizing the need for measuring P99 latency per turn, using prefetching, and budgeting memory tokens to reduce latency and improve user experience.
Introduces Human Bench 3.0, a benchmark for evaluating AI systems in comparison to human performance, with Humanity-2 AI-0 likely indicating a specific score or version.
A showcase of using the Qwen 3.8 27B model with 128k context to create a fully playable GTA Vice City-style game, including performance benchmarks and configuration tips.
GLM 5.3 benchmark results on SlopCodeBench show it scoring 47.1% on one subset and tying with Fable 5 and GPT-5.6 Sol on another, demonstrating that all models struggle with the unsaturated coding benchmark.
Elon Musk highlights continuous improvement in Intelligence/Joule, with Amjad Masad reporting an 18x increase in intelligence per joule over 16 months.
The article discusses the evaluation of the abliterated Qwen3.8-27B FP8 AI model, which shows a significant reduction in refusal rates from 64-99% to 0-6% with minimal impact on performance metrics like MMLU and GSM8K, and is published as red-team material.
A blog post illustrating how relying solely on the mean can be misleading when evaluating performance improvements, using synthetic latency data to show the importance of looking at the full distribution via percentiles, density plots, and CDFs.
This paper presents a complementary evaluation of PlanGPT, a large language model for automated planning, using plan cost and plan generation time metrics, and finds that PlanGPT performs no better than a greedy search strategy.
The author introduces the site plan for effectiveTPS, a tool designed to compare local AI models using a new 'effective TPS' metric alongside raw speed and latency. It aims to provide a simple leaderboard that highlights useful output quality over raw marketing numbers.