Tag
Proposes AdaSR, a framework enabling reasoning models to process streaming inputs adaptively, and HRPO, a hierarchical reinforcement learning method to optimize thinking allocation for accuracy-efficiency trade-offs.