Tag
iFixAi is an independent auditor that helps companies assess the trustworthiness of their AI agents through comprehensive audits covering AI misalignment.
Miles_Brundage advocates for mandatory auditing and enforceable safety standards for leading AI companies to improve oversight and security.
AVERI commends statements from Anthropic, OpenAI, and SpaceX on embedding independent experts in AI companies and commits to advancing standards for frontier AI auditing.
New research indicates that activation-based tools for AI auditing, such as activation oracles and SAEs, do not outperform simply reading transcripts, leading to a robust academic standard.
This study audits TikTok for harmful content exposure using multimodal language models across multiple countries and age groups, finding that keyword searches increase harmful content and that MLLMs provide a scalable tool for youth-safety audits.
This paper introduces DA-RAC, a distance-aware calibration method for LLM judges to enhance trustworthiness in AI auditing by using similar labeled anchors to reduce miscalibration and false-pass risks.
This paper argues that Latin America lacks a benchmark layer for native AI development and proposes an open, task-first EvalsHub infrastructure, with LatamBoard as its first regional instance, to audit AI systems and direct optimization toward local needs.
Miles Brundage discusses his move to focus on frontier AI auditing after leaving OpenAI, emphasizing that independent auditors can reassure AI companies that their peers are taking costly safety steps. AVERI supports this, advocating for paced development backed by independent oversight.
This paper introduces the nonuniformity principle for optimal human oversight placement in long AI workflows, demonstrating that oversight stages should be scheduled with non-decreasing gaps. The principle is validated empirically in literature review and website construction tasks.
Illinois enacts the strongest AI safety and accountability bill in the US, establishing the first frontier AI auditing requirement in the country.
This paper investigates how contextual framing affects LLM responses in mental health interactions, finding systematic behavioral variation and demonstrating that internal representations encode framing information throughout transformer layers.
A reflection on how LLM-based support automation leads to trust issues when errors occur, emphasizing the need for verification and auditability over pure accuracy improvement.
This paper introduces a statistical framework for adaptively auditing AI systems using Safe Anytime-Valid Inference (SAVI) to draw rigorous conclusions with limited data. It proposes a 'testing by betting' approach to validate model robustness while controlling type-I errors during adaptive sampling.