Tag
This paper introduces a safety cage framework for bounding the operational range of machine learning models in transit spectroscopy, demonstrating error reductions through indicator fusion to enhance reliability in safety-critical space missions.
This survey paper reviews 38 studies on safe LLM agents, highlighting key challenges such as specification translation bottlenecks, incomplete safety guarantees from enforcement methods like runtime monitoring, and the verifier tax that impedes safe task completion.
DreamGuard is a proactive runtime guardrail for LLM agents that uses a risk-aware world model to track latent state and predict future risks, enabling interventions before unsafe actions execute. It outperforms baselines on benchmarks and online evaluation with 25ms latency.
This survey synthesizes 257 papers on validating agentic AI systems, proposing a five-dimension taxonomy covering behavioral, safety, temporal, regulatory, and multi-agent concerns. It identifies gaps in temporal validity, runtime evidence maintenance, regulatory legibility, and open-ended multi-agent assurance, arguing that trustworthy deployment requires validating trajectories in context.
LabGuard introduces a framework that translates natural-language laboratory safety rules into executable runtime monitors for embodied agents, achieving a reduction in unsafe events from 39.5% to 23.8% while maintaining task success.
This paper studies how reasoning exchange among multiple AI agents can improve accuracy but also risk error propagation, proposing a runtime monitoring framework to prevent such propagation.
Hide-and-Seek is a framework that detects robot execution failures in VLA models by localizing failure-indicative actions through contrastive learning without step-level annotations, achieving state-of-the-art multi-task failure detection.
Proposes CPSS, a runtime safety mechanism that converts cumulative cost constraints into adaptive state-level thresholds for safe reinforcement learning in nonstationary environments, demonstrating reduced violations in highway merging scenarios.
This paper proposes reusable certified runtime monitors for past-time signal temporal logic (ptSTL) that use semantic latent representations to evaluate varying specifications without retraining, validated on pedestrian-crossroad and Waymo driving data.
This paper proposes Embedding Temporal Logic (ETL), a temporal logic that monitors perception-based autonomous systems directly in learned embedding spaces, enabling specification of high-level perceptual concepts and achieving strong empirical agreement with ground-truth semantics.