Tag
This paper presents a pre-registered audit of whether LLM-judged helpfulness can reliably distinguish answer-giving from pedagogical guidance in AI tutors. The authors find that general-purpose helpfulness is not a dependable pedagogy signal, recommending pedagogy-targeted rubrics and deterministic process measures instead.
Kimi K3 outperformed Fable/Opus 4.8 and GPT-5.6 Sol by finding 5 real bugs in a post-quantum cryptography project audit.
Using AI audit agents, zkSecurity discovered seven real bugs in Cloudflare's CIRCL cryptography library, including critical precision loss and access-control break. All bugs have been fixed upstream.
Code Shit Mountain Analysis Skills is an AI skill/prompt framework for generating rigorous, professional code review reports, supporting multiple audit modes and HTML output.
This article argues that AI agents making business recommendations must maintain complete audit logs to ensure trust and accountability for users, merchants, developers, and platforms, drawing parallels with traditional advertising systems.
A study analyzing 25,500 LLM resume evaluations across 10 models found a 45% bias rate driven by 'silent bias', with models inventing professional-sounding excuses to penalize candidates. It highlights significant variability in fairness and stability, with Claude, Mistral-Large, and Llama 4 being most stable, while Qwen and older Gemini models were volatile.
An AI-assisted security audit of FreeBSD uncovered 15 kernel vulnerabilities, including privilege escalations and a VM escape, and details the collaborative process of reporting and patching bugs with the FreeBSD team.