capability-profile

Tag

Cards List
#capability-profile

GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models

arXiv cs.AI ↗ · 2026-05-25 Cached

This paper introduces GENSTRAT, a benchmark that uses procedurally generated strategic environments to evaluate LLMs' strategic reasoning across multiple axes, addressing limitations of fixed game suites.

0 favorites 0 likes
← Back to home

Submit Feedback