Tag
This paper introduces a method for safe Bayesian optimization when safety is defined relative to a counterfactual baseline policy. It uses conformal prediction to estimate counterfactual outcomes and provides safety guarantees with user-specified violation rates.
SkillHarness is a framework that enables computer-use agents to safely learn and execute skills in dynamic environments by incorporating safety constraints and adaptive skill selection mechanisms, reducing unsafe rates by 57.1%.