Tag
This paper introduces a Bayesian framework for hierarchical LLM agents to decide when to escalate to stronger models during reasoning, formulating it as an optimal-stopping problem and providing theoretical guarantees on performance.
This paper introduces LLM-OSDA, a dynamic cost-per-click auction for native advertising in multi-turn LLM conversations, integrating Bellman optimal stopping, winner allocation, and envelope pricing. Experiments show an 11% net revenue improvement over fixed-timing baselines while maintaining user retention.
This paper introduces SARA, a sequential adaptive rollout allocation method for RLVR that abandons saturated groups early and reallocates the budget, achieving comparable accuracy with 22% fewer rollouts than dynamic sampling and up to 67% savings when combined.
This paper proposes CC-AOS, a structured amortized solver for finite-horizon optimal stopping problems that handles varying costs and horizons without retraining. It incorporates theoretical properties into the model architecture and demonstrates improved performance on benchmark tasks.
This paper formalizes the problem of when to invoke LLMs in streaming inference systems as a risk-based sequential stopping problem. It proves theoretical guarantees and empirically validates the framework on turbofan degradation data.
This paper introduces CARLOS, a deep reinforcement learning algorithm that learns continuous-time optimal stopping rules for American-style options using an aggregate deep neural network, effectively closing the Bermudan-American value gap with high computational efficiency.