Tag
The paper introduces LLAMIA-Bench, a benchmark for collaborative chess tasks between language models and non-language agents, and proposes latent state internalization to outperform text-based verbalization, with a 14B model matching or exceeding frontier models like GPT-5.1.