Tag
Introduces ClawArena-Team, a benchmark to measure the management ability of a single language model acting as a leader that creates, delegates to, and orchestrates subagents via dynamic workflows. Experiments reveal that privilege granting is a bottleneck, cost and management quality are decoupled, and most models cluster in performance while orchestration behaviors vary widely.