auditing

Tag

Cards List
#auditing

Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models

arXiv cs.CL ↗ · 2026-07-24 Cached

This paper investigates the mismatch between evaluation and deployment behavior in fine-tuned language models. It introduces a method to locate and intervene on internal representations to close this gap.

0 favorites 0 likes
#auditing

Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat

arXiv cs.CL ↗ · 2026-07-21 Cached

This paper develops the QQ equality from cognitive science into an audit criterion for LLMs, characterizes theoretical mechanism classes, and empirically tests on an open-weight instruction-tuned model, finding that saturation (near-deterministic responses) prevents distribution-level audits.

0 favorites 0 likes
#auditing

Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thought via Predicate Substitution

arXiv cs.AI ↗ · 2026-07-16 Cached

Introduces interventional grounding audits as a black-box, step-level test to check whether LLM chain-of-thought reasoning genuinely depends on its stated premises. Evaluated on ProntoQA with GPT-4o, achieving F1=0.806 on detecting proof-tree dependencies, significantly outperforming a self-consistency baseline.

0 favorites 0 likes
#auditing

ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI

arXiv cs.AI ↗ · 2026-07-13 Cached

ConceptSMILE is a perturbation-based auditing framework for evaluating the reliability of concept-based explainable AI, tested on retinal fundus images.

0 favorites 0 likes
#auditing

Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets

arXiv cs.AI ↗ · 2026-07-10 Cached

Introduces 'overthinking', a technique that amplifies reasoning weights from reasoning-distilled models to induce disclosure of hidden information in language models, demonstrating up to 10x greater secret leakage across 2B-32B models.

0 favorites 0 likes
#auditing

When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals

arXiv cs.AI ↗ · 2026-07-10 Cached

This paper audits whether self-consistency and cross-model agreement are reliable indicators of correctness in LLMs, finding that agreement is a weak, regime-dependent proxy and that frontier models exhibit overconfidence.

0 favorites 0 likes
#auditing

Adversarial Social Epistemology for Assemblies of Humans and Large Language Models

arXiv cs.AI ↗ · 2026-07-10 Cached

This paper proposes an adversarial social epistemology framework for analyzing trust, deception, and inference chains in communicative landscapes involving humans and large language models, and outlines mechanisms for auditing trust breaches.

0 favorites 0 likes
#auditing

Reasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations

arXiv cs.AI ↗ · 2026-07-09 Cached

This paper introduces reasoning consistency scanning, a method to audit whether chain-of-thought reasoning is logically consistent with the final answer in AI safety evaluations, distinguishing it from faithfulness. The authors formalize inconsistency subtypes, build a benchmark, implement a scanner, and report findings across models and tasks.

0 favorites 0 likes
#auditing

@OpenAI: To audit SWE-Bench Pro, we used model-based investigator agents alongside independent reviews from five independent exp…

X AI KOLs ↗ · 2026-07-08 Cached

OpenAI describes its audit of SWE-Bench Pro using model-based investigator agents and independent reviews from experienced software engineers to ensure thorough evaluation at scale.

0 favorites 0 likes
#auditing

Auditing of Unlearning Algorithms

arXiv cs.LG ↗ · 2026-07-08 Cached

Proposes a practical auditor that uses membership inference attacks to compute data-dependent lower bounds on the unlearning parameter, finding a sharp separation between certified algorithms (e.g., model clipping, rewind-to-delete) that achieve tight bounds and empirical methods (e.g., Hessian-based unlearning, gradient ascent) that exhibit large bounds, indicating poor unlearning.

0 favorites 0 likes
#auditing

When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions

arXiv cs.AI ↗ · 2026-07-07 Cached

This paper studies the evaluation of agentic AI systems that repair decision policies when per-state expert action labels are unavailable, using a hotel-pricing simulator with region-level diagnostic feedback. It finds that aggregate alignment can be misleading and proposes evaluating policy repair by closed-loop outcome rather than behavioral distance.

0 favorites 0 likes
#auditing

Auditing Forgetting in Limited Memory Language Models

arXiv cs.CL ↗ · 2026-07-02 Cached

This paper proposes a causal auditing framework to evaluate forgetting in Limited Memory Language Models by varying the database state during inference, discovering that parametric leakage is negligible and post-deletion correctness primarily arises from retrieval artifacts rather than residual parametric memory.

0 favorites 0 likes
#auditing

SentryCode: Real-time Auditor + Honeytokens for AI Coding Agents [P]

Reddit r/MachineLearning ↗ · 2026-07-02

SentryCode is an open-source kernel-level behavior auditing tool for AI coding agents that logs file/network/cue activity, uses honeypot tokens for zero-false-positive data breach detection, detects steganographic covert channels, and enforces policies, all running locally without network calls.

0 favorites 0 likes
#auditing

The Two Genie Game: Adoption and Welfare in Audit-Grounded AI Governance

arXiv cs.AI ↗ · 2026-06-30 Cached

This paper uses evolutionary game theory to model competition between a harm-minimizing AI agent and an approval-seeking (RLHF) agent in a community, analyzing conditions for adoption and welfare outcomes. The results show that while a self-audited agent can fixate, it is not sufficient to prevent community harm, and alignment and timeframe are critical.

0 favorites 0 likes
#auditing

@Miles_Brundage: I think we need federal AI regulation ASAP - something roughly along the lines of the Obernolte-Trahan but not blocking…

X AI KOLs Following ↗ · 2026-06-25 Cached

Miles Brundage calls for federal AI regulation with transparency and auditing requirements, noting that being pro-regulation helped a candidate in a primary.

0 favorites 0 likes
#auditing

@Miles_Brundage: Google just published an updated AI policy framework which articulates stronger and more detailed positions in some are…

X AI KOLs Following ↗ · 2026-06-25 Cached

Google published an updated AI policy framework with stronger and more detailed positions on auditing and other areas, marking a notable shift in their public stance.

0 favorites 0 likes
#auditing

Natural Identifiers for Privacy and Data Audits in Large Language Models

arXiv cs.LG ↗ · 2026-06-24 Cached

This paper introduces natural identifiers (NIDs) for post-hoc privacy auditing and dataset inference in large language models, eliminating the need for retraining or held-out datasets.

0 favorites 0 likes
#auditing

Do LLM Attribution Metrics Transfer? Auditing Retrieval-Augmented Generation Evaluation Across Datasets and Constructs

arXiv cs.CL ↗ · 2026-06-24 Cached

This paper audits eight automatic attribution metrics across three evaluation constructs for RAG systems, finding that no single metric transfers across datasets within the same construct, challenging the common practice of treating them as interchangeable.

0 favorites 0 likes
#auditing

Best tools for monitoring and auditing autonomous AI agent behavior at runtime, what's actually working in prod?

Reddit r/AI_Agents ↗ · 2026-06-23

A practitioner shares challenges and tools for monitoring autonomous AI agents in production, covering runtime prompt injection detection, tool-call auditing with reasoning traces, behavioral drift detection, and multi-agent authorization, while testing tools like Arize Phoenix, Protect AI Guardian, Metoro, Alice, Asqav, and Microsoft Agent Governance Toolkit.

0 favorites 0 likes
#auditing

ReasoningLens: Hierarchical Visualization and Diagnostic Auditing for Large Reasoning Models

Hugging Face Daily Papers ↗ · 2026-06-22 Cached

ReasoningLens is an open-source framework that provides hierarchical visualization and diagnostic auditing for complex reasoning chains in large reasoning models, enabling structured analysis and error detection.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback