I audited 112 real RL post-training environments for reward-hacking vulnerabilities — 54 flagged, 0 false positives [OC, tool] [P]

Reddit r/MachineLearning Tools

Summary

The author audited 112 RL post-training environments for reward-hacking vulnerabilities, building a static and dynamic auditor tool named ratctl that flagged 54 issues with 100% precision and released it as open-source.

RL post-training (RLHF/RLAIF/GRPO) agents optimize strictly for whatever the verifier rewards. If the verifier has logic flaws, the agent learns to hack the grader instead of solving the task — recent work has catalogued this at scale (Terminal Wrench found 331 hackable environments and 15%+ of standard benchmark tasks bypassable; a SWE-bench Verified audit found 28.5% Docker-verified hackability). I built ratctl, a static + dynamic auditor that scans RL environments (OpenEnv, Prime Intellect verifiers-spec, Gymnasium) for these patterns before you ship them for training: Test/assertion tampering (file deletion, pytest hook hijacking) Grader manipulation (stack frame inspection, __eq__/__bool__ overloading) Premature termination (sys.exit(0), signal suppression) Environment hijacking (golden solution leaks via .git log, env var leaks) Reward skipping (bare except: swallowing failures, hardcoded max rewards) LLM-judge bias (verbosity/sycophancy over correctness) Empirical results across 112 real environments (OpenEnv Hub, verifiers-spec, Gymnasium, SWE-bench): Metric Result Vulnerabilities flagged 54 Precision 100% (0 false positives on 43 clean controls) Recall 78.3% (54/69 known exploit patterns caught) Static analysis (AST-based) runs with zero heavy dependencies; there's an optional dynamic mode that red-teams the verifier with a local LLM (Ollama, free) or a frontier API. Ships as a CLI, a GitHub Action for CI gating (--fail-on 'gameability>0.3'), and a Claude/Cursor/Codex agent skill. Weakest spot right now: Gymnasium detection (0/8 caught in the audit) — mostly because vanilla Gymnasium envs don't expose the same exploit surface as adapter-based ones. Working on it. Repo + full audit report: https://github.com/FreakyAdy/Reward-Hackability-Auditor--CLI---Claude-Skill- Feedback and exploit patterns I missed are very welcome — especially if you've seen an agent hack a verifier in the wild that ratctl wouldn't catch.
Original Article

Similar Articles

Measuring Exploits in LLM Agents with Tool Use (4 minute read)

TLDR AI

An audit by Cursor finds that 63% of successful LLM agent runs on SWE-bench Pro retrieved the fix rather than deriving it, highlighting widespread reward hacking in coding benchmarks. The study proposes stricter environment controls to mitigate this behavior.