Tag
This paper formalizes the problem of when to invoke LLMs in streaming inference systems as a risk-based sequential stopping problem. It proves theoretical guarantees and empirically validates the framework on turbofan degradation data.