Tag
CrucibleBench places language models in a persistent MUD environment to evaluate agent behavior over 50 turns with hidden social objectives. The proof-of-concept release with 13 models revealed that using an LLM judge component can reorder leaderboards significantly, highlighting the need for reporting ranking stability under judge ablation.