Tag
Simon Willison introduces smevals, a small eval suite from Prime Radiant for evaluating models, prompts, and harnesses, with commands to run evals, grade results, and serve static HTML reports.
A developer created a starter kit and evaluation suite for Vapi-based AI receptionists after discovering silent mis-booking errors, sharing the tool to help others catch similar issues.