Tag
The article explains that SOC 2 compliance does not require pull requests; Amp demonstrates alternative controls like restricted push access, signed commits, automated CI, and audit trails to achieve compliance, emphasizing risk-based approaches over standard processes.
This paper presents IntelliAudit, a retrieval-grounded multi-agent system that uses large language models to evaluate IT audit controls against evidence corpora, generating cited recommendations and remediation guidance. The authors instantiate it on ISO/IEC 27001 and find it useful for audit preparation while emphasizing the need for human oversight.
This paper audits privacy risk in English-source multilingual RAG across five query languages, testing whether non-English queries increase PII leakage. Using a Qwen2.5-7B pipeline with two-stage defenses, it finds English has the highest point-estimate leak rate under output-only filtering, with residual leaks on Arabic and Swahili when the input judge is added.
Introduces FairFund-Bench, a benchmark for evaluating distributive bias in LLM resource allocation, showing that audit format changes the direction and magnitude of bias, and that causal framing effects dominate demographic effects.
This paper performs a forensic reproducibility audit of a radiology vision-language model benchmark, finding divergences between the intended protocol and released artifacts that invalidate the original claims. The authors propose a benchmark contract to expose such failure classes.
This paper audits major LLM benchmarks (GPQA, MMLU-Pro, MMMU-Pro) and finds 4.5–11.8% of items have wrong answer keys or are malformed, releasing corrected 'Clean' versions for drop-in use.
Bastillion 5.1 is a web-based SSH gateway and key management tool that now supports session auditing and replay, allowing users to record every SSH session and replay it on demand for compliance.
A tutorial on auditing your site's HTML for AI extractability issues, including building a Python audit script and adding a CI check to maintain extractability.
This paper audits five diversity measures for LLM ensembles, finding that their associations with majority-vote gain are heavily entangled with model capability and are unstable after controlling for capability. The only robust signal is a modest residual pairwise co-failure association.
California’s privacy agency has launched its first audit, targeting delivery and ride-share apps to ensure compliance with state privacy laws, focusing on gig economy platforms that collect extensive personal data on users and workers.
Apple has released System and Organization Controls (SOC) 3 audit reports for its Private Cloud Compute (PCC) Provisioning System, providing independent examinations of its controls for security, processing integrity, and confidentiality across multiple quarterly periods.
This paper evaluates whether multilingual sentence embeddings can replace translation for linguistic-integrated reliability auditing across multiple languages in educational assessments, finding that native-language embeddings reproduce translation-based reliability estimates closely.
This paper audits six KV-cache compression methods under query-agnostic protocols, finding that rankings change dramatically compared to query-aware evaluations, with implications for cache reuse in long-context inference.
A practitioner shares concerns about an upcoming audit revealing undocumented AI agents in production, highlighting governance gaps and risks with customer PII access.
A reflective article on the challenges of building AI agents for small businesses, comparing the process to pouring water into a leaky cup, where errors and bugs cause inefficiencies, but systematic debugging and rule-writing gradually improve the system.
This paper discusses testing methodologies for financial controls in ERP data provisioning, likely aimed at improving audit and compliance processes.
AuditWeave is a lightweight Python library that records workflow steps into a tamper-evident, hash-chained ledger for auditing AI-assisted and data-transformation workflows, enabling traceability and integrity verification.
Allstate accuses Broadcom of launching retaliatory audits after it terminated its VMware and CA contracts, highlighting ongoing legal battles between Broadcom and disgruntled enterprise customers.
A tweet thread shares a comprehensive prompt for auditing vibe-coded projects before launch, covering security, reliability, concurrency, accessibility, and UI consistency.
An auditor finds that 0 of 32 published "discoveries" by an autonomous research agent were novel as framed, with the dominant failure being over-labeling rather than bad measurement.