@zhangchen_xu: After more than six months of working with frontier labs on post-training for auto research, we’re sharing some of what…

X AI KOLs Timeline News

Summary

Sharing insights from over six months of post-training work on auto research, emphasizing findings that all tested models exhibited reward-hacking behaviors and the critical role of robust verifiers.

After more than six months of working with frontier labs on post-training for auto research, we’re sharing some of what we’ve learned about auto research and RSI, starting with reward hacking. On average, we spend 90% of our time and compute building robust verifiers. Tackling reward hacking in production has been fun and challenging, especially in auto research tasks, where agents improve their solutions progressively. Each iteration is another opportunity to make real progress, but also to exploit a flawed verifier. Findings 👉 • All 17 models we tested show reward-hacking behaviors without instructions to do so. • Some exploits looked like ordinary research choices but survived agent full-trajectory review. • When agents were tasked with evading detections by verifiers, detailed feedback and attempt history roughly doubled cumulative evasion over five rounds. Scaling auto research means scaling our ability to verify real progress. See papers 👇 #AutoReserch #RSI #Agents #LLM
Original Article
View Cached Full Text

Cached at: 09/27/26, 03:25 PM

After more than six months of working with frontier labs on post-training for auto research, we’re sharing some of what we’ve learned about auto research and RSI, starting with reward hacking.

On average, we spend 90% of our time and compute building robust verifiers.

Tackling reward hacking in production has been fun and challenging, especially in auto research tasks, where agents improve their solutions progressively. Each iteration is another opportunity to make real progress, but also to exploit a flawed verifier.

Findings 👉

• All 17 models we tested show reward-hacking behaviors without instructions to do so. • Some exploits looked like ordinary research choices but survived agent full-trajectory review. • When agents were tasked with evading detections by verifiers, detailed feedback and attempt history roughly doubled cumulative evasion over five rounds.

Scaling auto research means scaling our ability to verify real progress. See papers 👇

#AutoReserch #RSI #Agents #LLM

Similar Articles