tactical-reasoning

Tag

Cards List
#tactical-reasoning

DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat

arXiv cs.AI · 2026-08-03 Cached

DungeonBench is a new benchmark for evaluating tactical reasoning in Dungeons & Dragons combat, testing AI policies on rules-rich decision-making across single encounters and linked adventuring days. Frontier language models often win direct fights but struggle with resource budgeting and rest timing over longer horizons.

0 favorites 0 likes
#tactical-reasoning

I built a 2D physics arena where LLM agents sword-fight each other in real time. Turns out it's a surprisingly sharp test of tactical reasoning.

Reddit r/AI_Agents · 2026-06-15

Stickblade Arena is a new benchmark where LLM agents control ragdolls in a 2D physics sword-fighting simulator, testing multi-turn tactical reasoning, spatial awareness, and real-time decision-making under adversarial pressure. Early results reveal capability gaps: DeepSeek R1 dominates melee but fails at bow due to time limits, and small models excel at close-range fighting.

0 favorites 0 likes
← Back to home

Submit Feedback