Tag
This paper proposes predicting task difficulty for LLM agents without running expensive rollouts, studying the problem across 17 agentic benchmarks and showing that token-level entropy is a useful predictive signal.
Introduces Epi2Diff, a framework that maps LLM reasoning traces into cognitive episodes to predict human item difficulty, outperforming baselines and providing interpretable process evidence.