Tag
CogArena introduces a procedurally generated 13-paradigm benchmark to evaluate whether LLMs exhibit separable cognitive abilities or a single general competence, finding only weak support for stable five-dimensional profiles across 55 models.