Tag
The article proposes AdaThinking-E, a reinforcement learning framework that uses one-token entropy regulation to enable adaptive thinking in multimodal large language models, improving accuracy on complex tasks and efficiency on simple ones.
EntroRouter proposes a single-round model routing framework that uses entropy regulation to balance accuracy and computational cost, achieving 98.3% of the strongest expert's accuracy while reducing costs by 48.25%.