Tag
AgentAudit is an open, extensible framework for evaluating the full lifecycle of AI agents across capability, grounding, security, and behavioral dimensions, enabling precise failure attribution and highlighting trustworthiness differences among various language models.
The author adapts classical Islamic hadith verification methods to create a trust framework for multi-agent AI systems, releasing it as a paper and Python package (isnad).