Tag
This paper investigates the mismatch between evaluation and deployment behavior in fine-tuned language models. It introduces a method to locate and intervene on internal representations to close this gap.
This paper develops the QQ equality from cognitive science into an audit criterion for LLMs, characterizes theoretical mechanism classes, and empirically tests on an open-weight instruction-tuned model, finding that saturation (near-deterministic responses) prevents distribution-level audits.
Introduces interventional grounding audits as a black-box, step-level test to check whether LLM chain-of-thought reasoning genuinely depends on its stated premises. Evaluated on ProntoQA with GPT-4o, achieving F1=0.806 on detecting proof-tree dependencies, significantly outperforming a self-consistency baseline.
ConceptSMILE is a perturbation-based auditing framework for evaluating the reliability of concept-based explainable AI, tested on retinal fundus images.
Introduces 'overthinking', a technique that amplifies reasoning weights from reasoning-distilled models to induce disclosure of hidden information in language models, demonstrating up to 10x greater secret leakage across 2B-32B models.
This paper audits whether self-consistency and cross-model agreement are reliable indicators of correctness in LLMs, finding that agreement is a weak, regime-dependent proxy and that frontier models exhibit overconfidence.
This paper proposes an adversarial social epistemology framework for analyzing trust, deception, and inference chains in communicative landscapes involving humans and large language models, and outlines mechanisms for auditing trust breaches.
This paper introduces reasoning consistency scanning, a method to audit whether chain-of-thought reasoning is logically consistent with the final answer in AI safety evaluations, distinguishing it from faithfulness. The authors formalize inconsistency subtypes, build a benchmark, implement a scanner, and report findings across models and tasks.
OpenAI describes its audit of SWE-Bench Pro using model-based investigator agents and independent reviews from experienced software engineers to ensure thorough evaluation at scale.
Proposes a practical auditor that uses membership inference attacks to compute data-dependent lower bounds on the unlearning parameter, finding a sharp separation between certified algorithms (e.g., model clipping, rewind-to-delete) that achieve tight bounds and empirical methods (e.g., Hessian-based unlearning, gradient ascent) that exhibit large bounds, indicating poor unlearning.
This paper studies the evaluation of agentic AI systems that repair decision policies when per-state expert action labels are unavailable, using a hotel-pricing simulator with region-level diagnostic feedback. It finds that aggregate alignment can be misleading and proposes evaluating policy repair by closed-loop outcome rather than behavioral distance.
This paper proposes a causal auditing framework to evaluate forgetting in Limited Memory Language Models by varying the database state during inference, discovering that parametric leakage is negligible and post-deletion correctness primarily arises from retrieval artifacts rather than residual parametric memory.
SentryCode is an open-source kernel-level behavior auditing tool for AI coding agents that logs file/network/cue activity, uses honeypot tokens for zero-false-positive data breach detection, detects steganographic covert channels, and enforces policies, all running locally without network calls.
This paper uses evolutionary game theory to model competition between a harm-minimizing AI agent and an approval-seeking (RLHF) agent in a community, analyzing conditions for adoption and welfare outcomes. The results show that while a self-audited agent can fixate, it is not sufficient to prevent community harm, and alignment and timeframe are critical.
Miles Brundage calls for federal AI regulation with transparency and auditing requirements, noting that being pro-regulation helped a candidate in a primary.
Google published an updated AI policy framework with stronger and more detailed positions on auditing and other areas, marking a notable shift in their public stance.
This paper introduces natural identifiers (NIDs) for post-hoc privacy auditing and dataset inference in large language models, eliminating the need for retraining or held-out datasets.
This paper audits eight automatic attribution metrics across three evaluation constructs for RAG systems, finding that no single metric transfers across datasets within the same construct, challenging the common practice of treating them as interchangeable.
A practitioner shares challenges and tools for monitoring autonomous AI agents in production, covering runtime prompt injection detection, tool-call auditing with reasoning traces, behavioral drift detection, and multi-agent authorization, while testing tools like Arize Phoenix, Protect AI Guardian, Metoro, Alice, Asqav, and Microsoft Agent Governance Toolkit.
ReasoningLens is an open-source framework that provides hierarchical visualization and diagnostic auditing for complex reasoning chains in large reasoning models, enabling structured analysis and error detection.