runtime-monitoring

Tag

Cards List
#runtime-monitoring

Operational Range Bounding in Spectroscopy: A Safety Cage Framework for Machine Learning Models

arXiv cs.LG · 16h ago Cached

This paper introduces a safety cage framework for bounding the operational range of machine learning models in transit spectroscopy, demonstrating error reductions through indicator fusion to enhance reliability in safety-critical space missions.

0 favorites 0 likes
#runtime-monitoring

Toward Safe LLM Agents: A Survey of Specification, Verification, and Enforcement

arXiv cs.AI · 2026-08-18 Cached

This survey paper reviews 38 studies on safe LLM agents, highlighting key challenges such as specification translation bottlenecks, incomplete safety guarantees from enforcement methods like runtime monitoring, and the verifier tax that impedes safe task completion.

0 favorites 0 likes
#runtime-monitoring

DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model

arXiv cs.AI · 2026-08-07 Cached

DreamGuard is a proactive runtime guardrail for LLM agents that uses a risk-aware world model to track latent state and predict future risks, enabling interventions before unsafe actions execute. It outperforms baselines on benchmarks and online evaluation with 25ms latency.

0 favorites 0 likes
#runtime-monitoring

Beyond Component Testing: Validating Agentic AI Systems

arXiv cs.AI · 2026-08-03 Cached

This survey synthesizes 257 papers on validating agentic AI systems, proposing a five-dimension taxonomy covering behavioral, safety, temporal, regulatory, and multi-agent concerns. It identifies gaps in temporal validity, runtime evidence maintenance, regulatory legibility, and open-ended multi-agent assurance, arguing that trustworthy deployment requires validating trajectories in context.

0 favorites 0 likes
#runtime-monitoring

LabGuard: Grounding Natural-Language Laboratory Rules into Runtime Guards for Embodied Laboratory Agents

arXiv cs.AI · 2026-07-01 Cached

LabGuard introduces a framework that translates natural-language laboratory safety rules into executable runtime monitors for embodied agents, achieving a reduction in unsafe events from 39.5% to 23.8% while maintaining task success.

0 favorites 0 likes
#runtime-monitoring

Preventing Error Propagation in Multi-Agent AI through Runtime Monitoring

arXiv cs.AI · 2026-06-30 Cached

This paper studies how reasoning exchange among multiple AI agents can improve accuracy but also risk error propagation, proposing a runtime monitoring framework to prevent such propagation.

0 favorites 0 likes
#runtime-monitoring

Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring

Hugging Face Daily Papers · 2026-05-29 Cached

Hide-and-Seek is a framework that detects robot execution failures in VLA models by localizing failure-indicative actions through contrastive learning without step-level annotations, achieving state-of-the-art multi-task failure detection.

0 favorites 0 likes
#runtime-monitoring

From Cumulative Constraints to Adaptive Runtime Safety Control for Nonstationary Reinforcement Learning

arXiv cs.LG · 2026-05-20

Proposes CPSS, a runtime safety mechanism that converts cumulative cost constraints into adaptive state-level thresholds for safe reinforcement learning in nonstationary environments, demonstrating reduced violations in highway merging scenarios.

0 favorites 0 likes
#runtime-monitoring

Vision-Based Runtime Monitoring under Varying Specifications using Semantic Latent Representations

arXiv cs.LG · 2026-05-15 Cached

This paper proposes reusable certified runtime monitors for past-time signal temporal logic (ptSTL) that use semantic latent representations to evaluate varying specifications without retraining, validated on pedestrian-crossroad and Waymo driving data.

0 favorites 0 likes
#runtime-monitoring

Runtime Monitoring of Perception-Based Autonomous Systems via Embedding Temporal Logic

arXiv cs.LG · 2026-05-14 Cached

This paper proposes Embedding Temporal Logic (ETL), a temporal logic that monitors perception-based autonomous systems directly in learned embedding spaces, enabling specification of high-level perceptual concepts and achieving strong empirical agreement with ground-truth semantics.

0 favorites 0 likes
← Back to home

Submit Feedback