Tag
This paper reveals that while large language models appear robust to task-irrelevant context at the aggregate level, their predictions can flip on individual examples, with performance degrading on some and improving on others, highlighting tail risks that aggregate accuracy conceals.
This paper proposes a diagnostic framework to separate preprocessing pipeline instability from measurement method instability in LLM-based stance analysis of public discourse, finding that cross-method disagreement is larger and more systematic than pipeline effects, and that aggregate metrics can mask these instabilities.
This paper investigates why multi-step tool-use reinforcement learning (RL) often collapses or yields limited gains, identifying probability spikes in control tokens as a key cause. It shows that interleaving supervised fine-tuning with RL improves stability and explores various supervisory signals to guide robust training.