Tag
This paper investigates how reinforcement learning with verifiable rewards (RLVR) narrows the solution space in LLM reasoning by analyzing where diversity is lost, finding it concentrated at the 'entrance' of trajectories. It demonstrates interventions to recover breadth without compromising accuracy.
This paper presents a conflict-free path-planning algorithm for en-route air traffic control, designed to be interpretable and computationally efficient for human operators. The algorithm integrates three conflict detection methods and achieves fast computation times, demonstrated on a real-world sector.