optimal-stopping

Tag

Cards List
#optimal-stopping

Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents

arXiv cs.LG · 3d ago Cached

This paper introduces a Bayesian framework for hierarchical LLM agents to decide when to escalate to stronger models during reasoning, formulating it as an optimal-stopping problem and providing theoretical guarantees on performance.

0 favorites 0 likes
#optimal-stopping

LLM-OSDA: An Optimal-Stopping Dynamic Auction for Native Advertising in Multi-Turn LLM Conversations

arXiv cs.CL · 2026-08-04 Cached

This paper introduces LLM-OSDA, a dynamic cost-per-click auction for native advertising in multi-turn LLM conversations, integrating Bellman optimal stopping, winner allocation, and envelope pricing. Experiments show an 11% net revenue improvement over fixed-timing baselines while maintaining user retention.

0 favorites 0 likes
#optimal-stopping

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR

arXiv cs.LG · 2026-07-30 Cached

This paper introduces SARA, a sequential adaptive rollout allocation method for RLVR that abandons saturated groups early and reallocates the budget, achieving comparable accuracy with 22% fewer rollouts than dynamic sampling and up to 67% savings when combined.

0 favorites 0 likes
#optimal-stopping

CC-AOS: Cost- and Horizon-Conditioned Amortized Backward Induction for Finite-Horizon Optimal Stopping

arXiv cs.LG · 2026-07-28 Cached

This paper proposes CC-AOS, a structured amortized solver for finite-horizon optimal stopping problems that handles varying costs and horizons without retraining. It incorporates theoretical properties into the model architecture and demonstrates improved performance on benchmark tasks.

0 favorites 0 likes
#optimal-stopping

Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems

arXiv cs.LG · 2026-07-16 Cached

This paper formalizes the problem of when to invoke LLMs in streaming inference systems as a risk-based sequential stopping problem. It proves theoretical guarantees and empirically validates the framework on turbofan degradation data.

0 favorites 0 likes
#optimal-stopping

Continuous-time Optimal Stopping through Deep Reinforcement Learning

arXiv cs.LG · 2026-06-17 Cached

This paper introduces CARLOS, a deep reinforcement learning algorithm that learns continuous-time optimal stopping rules for American-style options using an aggregate deep neural network, effectively closing the Bermudan-American value gap with high computational efficiency.

0 favorites 0 likes
← Back to home

Submit Feedback