Tag
This paper introduces a Bayesian framework for hierarchical LLM agents to decide when to escalate to stronger models during reasoning, formulating it as an optimal-stopping problem and providing theoretical guarantees on performance.
This research paper investigates how data perturbations impact the accuracy and robustness of model cascades for efficient AI inference, identifying key failure modes and stressing the need for evaluation under distribution shift.