Tag
A blog post illustrating how relying solely on the mean can be misleading when evaluating performance improvements, using synthetic latency data to show the importance of looking at the full distribution via percentiles, density plots, and CDFs.
This paper presents a complementary evaluation of PlanGPT, a large language model for automated planning, using plan cost and plan generation time metrics, and finds that PlanGPT performs no better than a greedy search strategy.
The author introduces the site plan for effectiveTPS, a tool designed to compare local AI models using a new 'effective TPS' metric alongside raw speed and latency. It aims to provide a simple leaderboard that highlights useful output quality over raw marketing numbers.