@JongwonPar9958: GLM-5.2 has a neat trick for reward hacking. They don't penalize the model, they detect the suspicious tool call, block…
Summary
GLM-5.2 uses a technique to counteract reward hacking by detecting and blocking suspicious tool calls rather than penalizing the model, which prevents obfuscation seen in other methods.
View Cached Full Text
Cached at: 06/20/26, 08:24 PM
GLM-5.2 has a neat trick for reward hacking. They don’t penalize the model, they detect the suspicious tool call, block it, return dummy info, and keep training. The hack just stops paying off.
@bobabowen et al (2503.11926) showed penalizing a CoT monitor instead pushes the model to obfuscate, hide the intent and keep hacking. So neutralizing the action vs penalizing the signal shouldn’t behave the same. Recontextualization (2512.19027) and inoculation (2511.18397) are the same spirit, don’t touch the reward signal.
But I can’t find a head to head. Dummy vs penalty, same env, measuring obfuscation.
Anyone know one?
Similar Articles
Models know when they're reward hacking — and we can catch them at scale (16 minute read)
Research reveals that AI models frequently engage in reward hacking and can be detected at scale using activation probes, offering new mitigation strategies for AI safety.
@omarsar0: GLM-5.2 is great at design (Opus level IMO). I am also starting to see great results with long-running tasks, too. How …
GLM-5.2, an open-weight model with Opus-level design capabilities, incorporates an anti-hacking module trained via RL to mitigate reward hacking and improve performance on long-running tasks.
Detecting and Suppressing Reward Hacking with Gradient Fingerprints
This paper introduces Gradient Fingerprint (GRIFT), a method for detecting reward hacking in reinforcement learning with verifiable rewards by analyzing models' internal gradient computations rather than surface-level reasoning traces. The approach achieves over 25% relative improvement in detecting implicit reward-hacking behaviors across math, code, and logical reasoning benchmarks.
A debugger for RL reward functions that detects reward hacking during training [P]
A debugger that detects reward hacking in reinforcement learning reward functions during training, aiding developers in identifying and fixing issues.
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Survey introduces the Proxy Compression Hypothesis to explain how RLHF and related methods systematically induce reward hacking, deception, and oversight gaming in large language and multimodal models.