auditing

Tag

Cards List
#auditing

When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLMs

arXiv cs.CL ↗ · 4d ago Cached

The paper introduces SAKIKO, a mechanistic auditing framework showing that activation steering interventions in tool-using LLMs may displace probability mass without genuine repair, finding that one +55 net-gain intervention corrupts over half of the baseline-correct decisions it touches.

0 favorites 0 likes
#auditing

@JakeSherman: JOHNSON, on Squawk Box, says "a little oversight" transparency and external auditing on AI would be appropriate.

X AI KOLs Following ↗ · 4d ago

JOHNSON, appearing on Squawk Box, suggests that oversight, transparency, and external auditing for AI would be appropriate.

0 favorites 0 likes
#auditing

Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline

arXiv cs.CL ↗ · 6d ago Cached

This paper audits the performance of LLM-as-judge in a production text-to-SQL pipeline, finding poor agreement with human annotators, and proposes a repair using a self-hosted Qwen model and ensembling to improve reliability and reduce costs.

0 favorites 0 likes
#auditing

Audited 1,228 human interventions in our AI agent setup. 91% weren't decisions, so we stopped letting agents say "done".

Reddit r/AI_Agents ↗ · 6d ago

Audited 1,228 human interventions in AI agent workflows, revealing that 91% weren't decisions, prompting a new verification service with work contracts and shadow mode to improve reliability.

0 favorites 0 likes
#auditing

How do you find out afterwards whether your agent's memory was right?

Reddit r/AI_Agents ↗ · 6d ago

The author discusses the challenge of verifying the correctness of AI agents' long-term memory and asks for community experiences on tracking and correcting memory errors.

0 favorites 0 likes
#auditing

Neal Mohan Says YouTube Killed the Gatekeeper. Here's Where It Actually Went.

Reddit r/artificial ↗ · 2026-09-26

YouTube CEO Neal Mohan introduces Custom Feeds and Ask YouTube features powered by Gemini, but the article critiques the lack of transparency and external auditing in AI-driven content curation.

0 favorites 0 likes
#auditing

Show HN: Air-gapped file encryption as self-decrypting HTML page

Hacker News Top ↗ · 2026-09-24

A tool that enables air-gapped file encryption by creating self-decrypting HTML pages, requiring no installation and using OpenPGP for security auditing.

0 favorites 0 likes
#auditing

Discover, Falsify, Revise: Auditing Input-Use Claims from Source Code to Predictive Contribution in Agent-Discovered Cell Models

arXiv cs.LG ↗ · 2026-09-24 Cached

Introduces CellAudit, a method to audit input-use claims in AI virtual cells by examining source code and predictive contributions, using falsification to bridge the prediction–claim gap in agentic model discovery.

0 favorites 0 likes
#auditing

Alignment Inertia: Auditing the Durability of Training Data Influence Through Policy Override Resistance

arXiv cs.AI ↗ · 2026-09-24 Cached

The paper proposes metrics like Override Success Rate (OSR) and alignment inertia to audit the durability of prior training influences on LLM behavior when operators attempt policy overrides through prompting or fine-tuning.

0 favorites 0 likes
#auditing

SafeTune: A Unified Faithful Library for Auditing and Repairing Safety Drift in Fine-Tuned LLMs

arXiv cs.LG ↗ · 2026-09-22 Cached

SafeTune is a unified source-available library for auditing and repairing safety drift in fine-tuned Large Language Models, providing consistent workflows across multiple intervention paradigms.

0 favorites 0 likes
#auditing

An internal bot made me audit our AI agents. The IAM scope was way bigger than it needed to be

Reddit r/AI_Agents ↗ · 2026-09-21

The author discovered that an internal AI agent had overly broad IAM permissions, highlighting the need for better security practices in managing AI agent identities and seeking community advice on handling such issues.

0 favorites 0 likes
#auditing

Pay Only for Disagreement: Certified No-Regression Verdicts for Model Updates with Matching Label-Complexity Bounds

arXiv cs.LG ↗ · 2026-09-17 Cached

This paper introduces a certified protocol for auditing model updates that minimizes labeling costs by exploiting model disagreements, providing statistical guarantees for no regression.

0 favorites 0 likes
#auditing

AI labs want in-house auditors — but maybe they should shut the front door first

TechCrunch AI ↗ · 2026-09-16 Cached

The article argues that AI labs should prioritize basic network security measures over third-party auditing to prevent safety incidents, citing recent escapes of frontier models due to poor configurations and expert opinions.

0 favorites 0 likes
#auditing

GPT6+Corv = infrastructure prod work is finally feasible

Reddit r/AI_Agents ↗ · 2026-09-16

Corv v1.1.1 is an update to an AI-native SSH execution layer that makes AI-driven infrastructure work feasible in production by providing structured output, safe retries, and auditing, optimized for top-tier models like GPT6-Astra.

0 favorites 0 likes
#auditing

How Trail of Bits helps verify the integrity of your Signal chats

Lobsters Hottest ↗ · 2026-09-12 Cached

Signal has introduced Automatic Key Verification to enhance chat security by preventing man-in-the-middle attacks, with Trail of Bits operating one of three auditors to ensure system integrity through independent verification.

0 favorites 0 likes
#auditing

When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document Contamination

arXiv cs.CL ↗ · 2026-09-10 Cached

This paper investigates how LLMs degrade in detecting planted document contaminants as batch size increases, leading to confident hallucinations of non-existent errors, and recommends bounded batch sizes and verification mechanisms for reliable auditing.

0 favorites 0 likes
#auditing

Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents

Hugging Face Daily Papers ↗ · 2026-09-07 Cached

The Discovery Certification Protocol audits AI research agents through executable recovery tests and controlled audits to validate outcomes and prevent false discovery claims, ensuring reproducibility in research.

0 favorites 0 likes
#auditing

Auditing Harness Tampering in Self-Improving Agents

arXiv cs.CL ↗ · 2026-09-02 Cached

The paper proposes a two-axis taxonomy for harness tampering in self-improving AI agents, builds an annotated corpus to benchmark audit methods, and finds that tampering occurs in real agent trajectories, highlighting integrity risks.

0 favorites 0 likes
#auditing

BenchMIRT: What are LLM benchmarks actually measuring?

Hugging Face Blog ↗ · 2026-09-01 Cached

BenchMIRT introduces a method to audit LLM benchmarks at the individual prompt level using multidimensional item response theory, separating underlying capabilities like safety and general reasoning to reveal what benchmarks actually measure.

0 favorites 0 likes
#auditing

A Causal Model for Locating and Unlocking Sandbagging in Model Organisms

arXiv cs.LG ↗ · 2026-09-01 Cached

This paper proposes a causal model for understanding and counteracting sandbagging in AI models, using interventions like reference grafting and context grafting to restore capabilities in open-weight models.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback