safety-auditing

Tag

Cards List
#safety-auditing

Automata from Agent Traces: Failure and Next-Step Prediction

arXiv cs.AI · 3d ago Cached

This paper proposes using finite-state machines derived from LLM agent traces to predict failures and next steps, enhancing safety auditing and runtime monitoring for agents.

0 favorites 0 likes
#safety-auditing

When Certificates Fail: A Unified Safety Framework for Embedded Neural Interface Models

arXiv cs.LG · 2026-07-09 Cached

This paper demonstrates that formal robustness certificates for embedded neural interface models can pass even when task accuracy collapses under adversarial attack, and proposes a unified empirical audit framework to address alignment failures between training objectives and operational user welfare.

0 favorites 0 likes
#safety-auditing

Auditing Agent Harness Safety

arXiv cs.CL · 2026-05-15 Cached

This paper proposes HarnessAudit, a framework for auditing LLM agent execution trajectories beyond final outputs, focusing on boundary compliance, execution fidelity, and system stability. It introduces HarnessAudit-Bench with 210 tasks across eight domains and evaluates ten harness configurations, finding that task completion misaligns with safe execution and violations accumulate with trajectory length.

0 favorites 0 likes
← Back to home

Submit Feedback