Tag
DungeonBench is a new benchmark for evaluating tactical reasoning in Dungeons & Dragons combat, testing AI policies on rules-rich decision-making across single encounters and linked adventuring days. Frontier language models often win direct fights but struggle with resource budgeting and rest timing over longer horizons.
Stickblade Arena is a new benchmark where LLM agents control ragdolls in a 2D physics sword-fighting simulator, testing multi-turn tactical reasoning, spatial awareness, and real-time decision-making under adversarial pressure. Early results reveal capability gaps: DeepSeek R1 dominates melee but fails at bow due to time limits, and small models excel at close-range fighting.