Tag
The article proposes AdaThinking-E, a reinforcement learning framework that uses one-token entropy regulation to enable adaptive thinking in multimodal large language models, improving accuracy on complex tasks and efficiency on simple ones.
AGORA is a new benchmark for evaluating large language models on archive-grounded reasoning tasks across workplace documents, comprising 362 questions over 9,664 real documents. The strongest model achieves only 59.4% accuracy, highlighting substantial room for improvement.