runtime-guardrails

Tag

Cards List
#runtime-guardrails

Safeguarding LLM Agents from Misalignment through Provenance Analysis

arXiv cs.CL · 2026-07-03 Cached

This paper proposes a provenance-based framework and multi-stage pipeline, \tool, to detect misalignment in LLM agents' tool invocations before execution, reducing error rates significantly compared to LLM-as-a-judge baselines.

0 favorites 0 likes
#runtime-guardrails

AI Agents are deleting DBs. Would you use a "Policy-as-Code" Gateway to stop them?

Reddit r/AI_Agents · 2026-06-22

The article highlights the risk of AI agents performing destructive actions like deleting databases and proposes a Runtime Policy Gateway that uses Policy-as-Code to intercept and block non-compliant agent actions in real time, asking if users would adopt such a security tool.

0 favorites 0 likes
#runtime-guardrails

Silent Failures in Physical AI: A Literature Review of Runtime Action Authorization for Autonomous Systems

Hugging Face Daily Papers · 2026-05-23 Cached

This literature review identifies and analyzes the problem of silent failures in physical AI systems, where black-box models may execute harmful actions without detection. It proposes a taxonomy of runtime guardrail functions and outlines evaluation requirements for safe autonomous systems.

0 favorites 0 likes
← Back to home

Submit Feedback