social-objectives

Tag

Cards List
#social-objectives

Can a MUD evaluate LLMs? A $99 proof of concept

Hacker News Top · 2026-07-22 Cached

CrucibleBench places language models in a persistent MUD environment to evaluate agent behavior over 50 turns with hidden social objectives. The proof-of-concept release with 13 models revealed that using an LLM judge component can reorder leaderboards significantly, highlighting the need for reporting ranking stability under judge ablation.

0 favorites 0 likes
← Back to home

Submit Feedback