Tag
This paper proposes a provenance-based framework and multi-stage pipeline, \tool, to detect misalignment in LLM agents' tool invocations before execution, reducing error rates significantly compared to LLM-as-a-judge baselines.
The article highlights the risk of AI agents performing destructive actions like deleting databases and proposes a Runtime Policy Gateway that uses Policy-as-Code to intercept and block non-compliant agent actions in real time, asking if users would adopt such a security tool.
This literature review identifies and analyzes the problem of silent failures in physical AI systems, where black-box models may execute harmful actions without detection. It proposes a taxonomy of runtime guardrail functions and outlines evaluation requirements for safe autonomous systems.