ai-audit

Tag

Cards List
#ai-audit

Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models

arXiv cs.CL · 3d ago Cached

This paper presents a pre-registered audit of whether LLM-judged helpfulness can reliably distinguish answer-giving from pedagogical guidance in AI tutors. The authors find that general-purpose helpfulness is not a dependable pedagogy signal, recommending pedagogy-targeted rubrics and deterministic process measures instead.

0 favorites 0 likes
#ai-audit

I gave Kimi K3 a shot at auditing my post-quantum crypto project, it found 5 real bugs Fable/Opus 4.8 and GPT-5.6 Sol had all missed

Reddit r/LocalLLaMA · 2026-07-20

Kimi K3 outperformed Fable/Opus 4.8 and GPT-5.6 Sol by finding 5 real bugs in a post-quantum cryptography project audit.

0 favorites 0 likes
#ai-audit

AI Meets Cryptography 1: What AI Found in Cloudflare's Circl

Hacker News Top · 2026-07-07 Cached

Using AI audit agents, zkSecurity discovered seven real bugs in Cloudflare's CIRCL cryptography library, including critical precision loss and access-control break. All bugs have been fixed upstream.

0 favorites 0 likes
#ai-audit

@geekbb: Code Shit Mountain Analysis Skills generates rigorous, professional code review reports. The name may be crude, but every report is calm, structured, evidence-driven, and actionable. https://github.com/XiNian-dada/Fuck_My_Shit_Mountain…

X AI KOLs Timeline · 2026-07-01 Cached

Code Shit Mountain Analysis Skills is an AI skill/prompt framework for generating rigorous, professional code review reports, supporting multiple audit modes and HTML output.

0 favorites 0 likes
#ai-audit

AI business may require complete audit logs.

Reddit r/AI_Agents · 2026-07-01

This article argues that AI agents making business recommendations must maintain complete audit logs to ensure trust and accountability for users, merchants, developers, and platforms, drawing parallels with traditional advertising systems.

0 favorites 0 likes
#ai-audit

I analyzed 25,500 LLM resume screenings to measure hiring bias. The results are a wake-up call.

Reddit r/artificial · 2026-06-01

A study analyzing 25,500 LLM resume evaluations across 10 models found a 45% bias rate driven by 'silent bias', with models inventing professional-sounding excuses to penalize candidates. It highlights significant variability in fairness and stability, with Claude, Mistral-Large, and Llama 4 being most stable, while Qwen and older Gemini models were volatile.

0 favorites 0 likes
#ai-audit

An AI audit of FreeBSD

Lobsters Hottest · 2026-05-29 Cached

An AI-assisted security audit of FreeBSD uncovered 15 kernel vulnerabilities, including privilege escalations and a VM escape, and details the collaborative process of reporting and patching bugs with the FreeBSD team.

0 favorites 0 likes
← Back to home

Submit Feedback